:TITLE: Regular Expressions 101 ;# ;# RCSID: $Header: /cvsroot/tcl/tcltutorial/original/Tcl20.lsn,v 1.1 2004/11/04 16:01:14 davidw Exp $ ;# Copyright (c) 1995 Clif Flynt ;# 9300 Fleming Rd. ;# Dexter, MI 48130 ;# clif@cflynt.com ;# See file "NOTICE" for licensing terms. ;# :LESSON_TEXT_START_LEVEL 0:
Tcl also supports string operations that use the regular expressions package written by Henry Spencer. Several commands can access these methods with a -regexp argument, see the man pages for which commands support regular expressions.
There are also two explicit commands for parsing regular expressions.
Regular expressions are similar to the globbing that was discussed in lessons 16 and 18. The main difference is in the way that sets of matched characters are handled. In globbing the only way to select sets of unknown text is the * symbol. This matches to any quantity of any character.
In regular expression parsing, the * symbol matches zero or more occurrences of the character immediately proceeding the *. For example a* would match a, aaaaa, or a blank string. If the character directly before the * is a set of characters within square brackets, then the * will match any quantity of all of these characters. For example, [a-c]* would match aa, abc, aabcabc, or again, an empty string.
The + symbol behaves roughly the same as the *, except that it requires at least one character to match. For example, [a-c]* would match aa, abc, or aabcabc, but not an empty string.
The regexpcommand is similar to the string match command in that it matches an exp against a string. It is different in that it can match a portion of a string, instead of the entire string, and will place the characters matched into the matchVar variable.
If a match is found to the portion of a regular expression enclosed within parentheses, regexp will copy the subset of matching characters is to the subSpec argument. This can be used to parse simple strings.
Regsub will copy the contents of the string to a new variable, substituting the characters that match exp with the characters in subSpec. If subSpec contains a & or \0, then those characters will be replaced by the characters that matched exp. If the number following a backslash is 1-9, then that backslash sequence will be replaced by the appropriate portion of exp that is enclosed within parentheses. :TEXT_END: :LESSON_TEXT_START_LEVEL 1: Tcl also supports string operations that use the regular expressions package written by Henry Spencer. Several commands can access these methods with a -regexp argument, see the man pages for which commands support regular expressions.
There are also two explicit commands for parsing regular expressions.
Regular expressions are similar to the globbing that was discussed in lessons 16 and 18. The main difference is in the way that sets of matched characters are handled. In globbing the only way to select sets of unknown text is the * symbol. This matches to any quantity of any character.
In regular expression parsing, the * symbol matches zero or more occurrences of the character immediately proceeding the *. For example a* would match a, aaaaa, or a blank string. If the character directly before the * is a set of characters within square brackets, then the * will match any quantity of all of these characters. For example, [a-c]* would match aa, abc, aabcabc, or again, an empty string.
The + symbol behaves roughly the same as the *, except that it requires at least one character to match. For example, [a-c]+ would match a, abc, or aabcabc, but not an empty string.
Regular expression parsing has a more powerful manner of parsing square brackets than globbing. With globbing you can use the square brackets to enclose a set of characters any of which will be a match. Regular expression parsing also includes a method of selecting any character not in a set. If the first character after the [ is a caret (^), then the regular expression parser will match any character not in the set of characters between the square brackets. A caret can be included in the set of characters to match (or not) by placing it in any position other than the first.
The regexpcommand is similar to the string match command in that it matches an exp against a string. It is different in that it can match a portion of a string, instead of the entire string, and will place the characters matched into the matchVar variable.
If a match is found to the portion of a regular expression enclosed within parentheses, regexp will copy the subset of matching characters is to the subSpec argument. This can be used to parse simple strings.
Regsub will copy the contents of the string to a new variable, substituting the characters that match exp with the characters in subSpec. If subSpec contains a & or \0, then those characters will be replaced by the characters that matched exp. If the number following a backslash is 1-9, then that backslash sequence will be replaced by the appropriate portion of exp that is enclosed within parentheses.
Note that the exp argument to regexp or regsub is
processed by the Tcl substitution pass, and be certain to escape characters
as required.
:TEXT_END:
:LESSON_TEXT_START_LEVEL 2:
Tcl also supports string operations that use the regular expressions package written by Henry Spencer. Several commands can access these methods with a -regexp argument, see the man pages for which commands support regular expressions.
There are also two explicit commands for parsing regular expressions.
Regular expressions are similar to the globbing that was discussed in lessons 16 and 18. The main difference is in the way that sets of matched characters are handled. In globbing the only way to select sets of unknown text is the * symbol. This matches to any quantity of any character.
In regular expression parsing, the * symbol matches zero or more occurrences of the character immediately proceeding the *. For example a* would match a, aaaaa, or a blank string. If the character directly before the * is a set of characters within square brackets, then the * will match any quantity of all of these characters. For example, [a-c]* would match aa, abc, aabcabc, or again, an empty string.
The + symbol behaves roughly the same as the *, except that it requires at least one character to match. For example, [a-c]+ would match a, abc, or aabcabc, but not an empty string.
Regular expression parsing has a more powerful manner of parsing square brackets than globbing. With globbing you can use the square brackets to enclose a set of characters any of which will be a match. Regular expression parsing also includes a method of selecting any character not in a set. If the first character after the [ is a caret (^), then the regular expression parser will match any character not in the set of characters between the square brackets. A caret can be included in the set of characters to match (or not) by placing it in any position other than the first.
The regexpcommand is similar to the string match command in that it matches an exp against a string. It is different in that it can match a portion of a string, instead of the entire string, and will place the characters matched into the matchVar variable.
If a match is found to the portion of a regular expression enclosed within parentheses, regexp will copy the subset of matching characters is to the subSpec argument. This can be used to parse simple strings.
Regsub will copy the contents of the string to a new variable, substituting the characters that match exp with the characters in subSpec. If subSpec contains a & or \0, then those characters will be replaced by the characters that matched exp. If the number following a backslash is 1-9, then that backslash sequence will be replaced by the appropriate portion of exp that is enclosed within parentheses.
If you haven't already done so, hit the run example button, and look at the output while you read the rest of this lesson.
The exp argument to regexp or regsub is
processed by the Tcl substitution pass. This affects the mechanics of how
regular expressions can be formed. If the regular expression has no
backslash sequences that you wish to have converted by the Tcl substitution
phase, then the exp argument can be enclosed in braces. If you wish
any substitutions to occur, however, the exp argument must be enclosed
within quotes. In this case, characters that have meaning to the Tcl
substitution phase must be escaped with a backslash.
In these examples, there are no backslash sequences, so either braces or quotes could be used to group the exp string. In the first two examples, the square bracket is used to define a range of acceptable characters, so it is simpler to use the brackets in these cases.
The first example simply matches the first set of lowercase characters. It won't match an upper case character, and it won't match a space. So, the match is "here", the first four lower case letters, bounded on one end by the upper case "W", and on the other end by the space.
The second example shows the first 2 words of the phrase being parsed
into separate variables. The sequence ([A-Za-z]+) selects
any number (but at least one) of upper or lowercase letters. The parentheses
around this sequence causes the characters that match this regular expression
to be copied into the first of the subSpec variables. In this case,
the letters "Where" are copied into sub1.
The next sequence is the " +". This matches any number of
spaces, but at least one space.
The last sequence in the exp is ([a-z]+), which matches
any quantity (but at least one) of lower case letters, and copied the
matching characters into the second subSpec variable.
The last example shows a trivial regular expression "way" being
replaced by the string "lawsuit" in the sample phrase.
Try changing the "way" to a generalized regular expression,
and see what happens.
:TEXT_END:
:CODE_START:
set sample "Where there is a will, There is a way."
set result [regexp {[a-z]+} $sample match]
puts "Result: $result match: $match"
set result [regexp {([A-Za-z]+) +([a-z]+)} $sample match sub1 sub2 ]
puts "Result: $result Match: $match 1: $sub1 2: $sub2"
regsub "way" $sample "lawsuit" sample2
puts "New: $sample2"
:TEXT_END: