Outils pour utilisateurs

Outils du site


francois:an_introduction_to_sed

Différences

Ci-dessous, les différences entre deux révisions de la page.

Lien vers cette vue comparative

francois:an_introduction_to_sed [2013/05/14 08:30] – créée francoisfrancois:an_introduction_to_sed [2013/05/15 10:24] (Version actuelle) – francois
Ligne 1: Ligne 1:
 ====== Introduction to Sed ====== ====== Introduction to Sed ======
- 
  
 How to use sed, a special editor for modifying files automatically. If you want to write a program to make changes in a file, sed is the tool to use. How to use sed, a special editor for modifying files automatically. If you want to write a program to make changes in a file, sed is the tool to use.
Ligne 14: Ligne 13:
 Anyhow, sed is a marvelous utility. Unfortunately, most people never learn its real power. The language is very simple, but the documentation is terrible. The Solaris on-line manual pages for sed are five pages long, and two of those pages describe the 34 different errors you can get. A program that spends as much space documenting the errors than it does documenting the language has a serious learning curve. Anyhow, sed is a marvelous utility. Unfortunately, most people never learn its real power. The language is very simple, but the documentation is terrible. The Solaris on-line manual pages for sed are five pages long, and two of those pages describe the 34 different errors you can get. A program that spends as much space documenting the errors than it does documenting the language has a serious learning curve.
  
-Do not fret! It is not your fault you don't understand sed. I will cover sed completely. But I will describe the features in the order that I learned them. I didn't learn everything at once. You don't need to either. +**Do not fret!** It is not your fault you don't understand sed. I will cover sed completely. But I will describe the features in the order that I learned them. I didn't learn everything at once. You don't need to either. 
-The essential command: s for substitution+ 
 +===== The essential command: s for substitution =====
  
 Sed has several commands, but most people only learn the substitute command: s. The substitute command changes all occurrences of the regular expression into a new value. A simple example is changing "day" in the "old" file to "night" in the "new" file: Sed has several commands, but most people only learn the substitute command: s. The substitute command changes all occurrences of the regular expression into a new value. A simple example is changing "day" in the "old" file to "night" in the "new" file:
  
-sed s/day/night/ <old >new+  sed s/day/night/ <old >new
  
 Or another way (for Unix beginners), Or another way (for Unix beginners),
  
-sed s/day/night/ old >new+  sed s/day/night/ old >new
  
 and for those who want to test this: and for those who want to test this:
  
-echo day | sed s/day/night/+  echo day | sed s/day/night/
  
 This will output "night". This will output "night".
Ligne 33: Ligne 33:
 I didn't put quotes around the argument because this example didn't need them. If you read my earlier tutorial, you would understand why it doesn't need quotes. However, I recommend you do use quotes. If you have meta-characters in the command, quotes are necessary. And if you aren't sure, it's a good habit, and I will henceforth quote future examples to emphasize the "best practice." Using the strong (single quote) character, that would be: I didn't put quotes around the argument because this example didn't need them. If you read my earlier tutorial, you would understand why it doesn't need quotes. However, I recommend you do use quotes. If you have meta-characters in the command, quotes are necessary. And if you aren't sure, it's a good habit, and I will henceforth quote future examples to emphasize the "best practice." Using the strong (single quote) character, that would be:
  
-sed 's/day/night/' <old >new+  sed 's/day/night/' <old >new
  
 I must emphasize the the sed editor changes exactly what you tell it to. So if you executed I must emphasize the the sed editor changes exactly what you tell it to. So if you executed
  
-echo Sunday | sed 's/day/night/' <old >new+  echo Sunday | sed 's/day/night/' <old >new
  
 This would output the word "Sunnight" bacause sed found the string "day" in the input. This would output the word "Sunnight" bacause sed found the string "day" in the input.
Ligne 43: Ligne 43:
 There are four parts to this substitute command: There are four parts to this substitute command:
  
-s Substitute command +  *   s Substitute command 
- +  *   /../../ Delimiter 
-/../../ Delimiter +  *   day Regular Expression Pattern Search Pattern 
- +  *   night Replacement string
-day Regular Expression Pattern Search Pattern +
- +
-night Replacement string+
  
 The search pattern is on the left hand side and the replacement string is on the right hand side. The search pattern is on the left hand side and the replacement string is on the right hand side.
  
 We've covered  quoting and  regular expressions.. That's 90% of the effort needed to learn the substitute command. To put it another way, you already know how to handle 90% of the most frequent uses of sed. There are a ... few fine points that an future sed expert should know about. (You just finished section 1. There's only 63 more sections to cover. :-) Oh. And you may want to bookmark this page, .... just in case you don't finish. We've covered  quoting and  regular expressions.. That's 90% of the effort needed to learn the substitute command. To put it another way, you already know how to handle 90% of the most frequent uses of sed. There are a ... few fine points that an future sed expert should know about. (You just finished section 1. There's only 63 more sections to cover. :-) Oh. And you may want to bookmark this page, .... just in case you don't finish.
-The slash as a delimiter+ 
 +===== The slash as a delimiter =====
  
 The character after the s is the delimiter. It is conventionally a slash, because this is what ed, more, and vi use. It can be anything you want, however. If you want to change a pathname that contains a slash - say /usr/local/bin to /common/bin - you could use the backslash to quote the slash: The character after the s is the delimiter. It is conventionally a slash, because this is what ed, more, and vi use. It can be anything you want, however. If you want to change a pathname that contains a slash - say /usr/local/bin to /common/bin - you could use the backslash to quote the slash:
  
-sed 's/\/usr\/local\/bin/\/common\/bin/' <old >new+  sed 's/\/usr\/local\/bin/\/common\/bin/' <old >new
  
 Gulp. Some call this a 'Picket Fence' and it's ugly. It is easier to read if you use an underline instead of a slash as a delimiter: Gulp. Some call this a 'Picket Fence' and it's ugly. It is easier to read if you use an underline instead of a slash as a delimiter:
  
-sed 's_/usr/local/bin_/common/bin_' <old >new+  sed 's_/usr/local/bin_/common/bin_' <old >new
  
 Some people use colons: Some people use colons:
  
-sed 's:/usr/local/bin:/common/bin:' <old >new+  sed 's:/usr/local/bin:/common/bin:' <old >new
  
 Others use the "|" character. Others use the "|" character.
  
-sed 's|/usr/local/bin|/common/bin|' <old >new+  sed 's|/usr/local/bin|/common/bin|' <old >new
  
 Pick one you like. As long as it's not in the string you are looking for, anything goes. And remember that you need three delimiters. If you get a "Unterminated `s' command" it's because you are missing one of them. Pick one you like. As long as it's not in the string you are looking for, anything goes. And remember that you need three delimiters. If you get a "Unterminated `s' command" it's because you are missing one of them.
-Using & as the matched string+ 
 +===== Using & as the matched string =====
  
 Sometimes you want to search for a pattern and add some characters, like parenthesis, around or near the pattern you found. It is easy to do this if you are looking for a particular string: Sometimes you want to search for a pattern and add some characters, like parenthesis, around or near the pattern you found. It is easy to do this if you are looking for a particular string:
  
-sed 's/abc/(abc)/' <old >new+  sed 's/abc/(abc)/' <old >new
  
 This won't work if you don't know exactly what you will find. How can you put the string you found in the replacement string if you don't know what it is? This won't work if you don't know exactly what you will find. How can you put the string you found in the replacement string if you don't know what it is?
Ligne 83: Ligne 82:
 The solution requires the special character "&." It corresponds to the pattern found. The solution requires the special character "&." It corresponds to the pattern found.
  
-sed 's/[a-z]*/(&)/' <old >new+  sed 's/[a-z]*/(&)/' <old >new
  
 You can have any number of "&" in the replacement string. You could also double a pattern, e.g. the first number of a line: You can have any number of "&" in the replacement string. You could also double a pattern, e.g. the first number of a line:
  
-% echo "123 abc" | sed 's/[0-9]*/& &/' +  % echo "123 abc" | sed 's/[0-9]*/& &/' 
- +  123 123 abc
-123 123 abc+
  
 Let me slightly amend this example. Sed will match the first string, and make it as greedy as possible. The first match for '[0-9]*' is the first character on the line, as this matches zero of more numbers. So if the input was "abc 123" the output would be unchanged (well, except for a space before the letters). A better way to duplicate the number is to make sure it matches a number: Let me slightly amend this example. Sed will match the first string, and make it as greedy as possible. The first match for '[0-9]*' is the first character on the line, as this matches zero of more numbers. So if the input was "abc 123" the output would be unchanged (well, except for a space before the letters). A better way to duplicate the number is to make sure it matches a number:
  
-% echo "123 abc" | sed 's/[0-9][0-9]*/& &/'+  % echo "123 abc" | sed 's/[0-9][0-9]*/& &/' 
 +  123 123 abc
  
-123 123 abc+The string "abc" is unchanged, because it was not matched by the regular expression. If you wanted to eliminate "abc" from the output, you must expand the the regular expression to match the rest of the line and explicitly exclude part of the expression using "(", ")" and "\1", which is the next topic.
  
-The string "abc" is unchanged, because it was not matched by the regular expression. If you wanted to eliminate "abc" from the output, you must expand the the regular expression to match the rest of the line and explicitly exclude part of the expression using "(", ")" and "\1", which is the next topic. +===== Using \1 to keep part of the pattern =====
-Using \1 to keep part of the pattern+
  
 I have already described the use of "(" ")" and "1" in my tutorial on  regular expressions. To review, the escaped parentheses (that is, parentheses with backslashes before them) remember portions of the regular expression. You can use this to exclude part of the regular expression. The "\1" is the first remembered pattern, and the "\2" is the second remembered pattern. Sed has up to nine remembered patterns. I have already described the use of "(" ")" and "1" in my tutorial on  regular expressions. To review, the escaped parentheses (that is, parentheses with backslashes before them) remember portions of the regular expression. You can use this to exclude part of the regular expression. The "\1" is the first remembered pattern, and the "\2" is the second remembered pattern. Sed has up to nine remembered patterns.
Ligne 104: Ligne 102:
 If you wanted to keep the first word of a line, and delete the rest of the line, mark the important part with the parenthesis: If you wanted to keep the first word of a line, and delete the rest of the line, mark the important part with the parenthesis:
  
-sed 's/\([a-z]*\).*/\1/'+  sed 's/\([a-z]*\).*/\1/'
  
 I should elaborate on this. Regular exprssions are greedy, and try to match as much as possible. "[a-z]*" matches zero or more lower case letters, and tries to be as big as possible. The ".*" matches zero or more characters after the first match. Since the first one grabs all of the lower case letters, the second matches anything else. Therefore if you type I should elaborate on this. Regular exprssions are greedy, and try to match as much as possible. "[a-z]*" matches zero or more lower case letters, and tries to be as big as possible. The ".*" matches zero or more characters after the first match. Since the first one grabs all of the lower case letters, the second matches anything else. Therefore if you type
  
-echo abcd123 | sed 's/\([a-z]*\).*/\1/'+  echo abcd123 | sed 's/\([a-z]*\).*/\1/'
  
 This will output "abcd" and delete the numbers. This will output "abcd" and delete the numbers.
Ligne 114: Ligne 112:
 If you want to switch two words around, you can remember two patterns and change the order around: If you want to switch two words around, you can remember two patterns and change the order around:
  
-sed 's/\([a-z]*\) \([a-z]*\)/\2 \1/'+  sed 's/\([a-z]*\) \([a-z]*\)/\2 \1/'
  
 Note the space between the two remembered patterns. This is used to make sure two words are found. Note the space between the two remembered patterns. This is used to make sure two words are found.
Ligne 120: Ligne 118:
 The "\1" doesn't have to be in the replacement string (in the right hand side). It can be in the pattern you are searching for (in the left hand side). If you want to eliminate duplicated words, you can try: The "\1" doesn't have to be in the replacement string (in the right hand side). It can be in the pattern you are searching for (in the left hand side). If you want to eliminate duplicated words, you can try:
  
-sed 's/\([a-z]*\) \1/\1/'+  sed 's/\([a-z]*\) \1/\1/'
  
 You can have up to nine values: "\1" thru "\9." You can have up to nine values: "\1" thru "\9."
-Substitute Flags+ 
 +===== Substitute Flags =====
  
 You can add additional flags after the last delimiter. These flags can specify what happens when there is more than one occurrence of a pattern on a single line, and what to do if a substitution is found. Let me describe them. You can add additional flags after the last delimiter. These flags can specify what happens when there is more than one occurrence of a pattern on a single line, and what to do if a substitution is found. Let me describe them.
-/g - Global replacement+ 
 +===== /g - Global replacement =====
  
 Most Unix utilties work on files, reading a line at a time. Sed, by default, is the same way. If you tell it to change a word, it will only change the first occurrence of the word on a line. You may want to make the change on every word on the line instead of the first. For an example, let's place parentheses around words on a line. Instead of using a pattern like "[A-Za-z]*" which won't match words like "won't," we will use a pattern, "[^ ]*," that matches everything except a space. Well, this will also match anything because "*" means zero or more. The current version of sed can get unhappy with patterns like this, and generate errors like "Output line too long" or even run forever. I consider this a bug, and have reported this to Sun. As a work-around, you must avoid matching the null string when using the "g" flag to sed. A work-around example is: "[^ ][^ ]*." The following will put parenthesis around the first word: Most Unix utilties work on files, reading a line at a time. Sed, by default, is the same way. If you tell it to change a word, it will only change the first occurrence of the word on a line. You may want to make the change on every word on the line instead of the first. For an example, let's place parentheses around words on a line. Instead of using a pattern like "[A-Za-z]*" which won't match words like "won't," we will use a pattern, "[^ ]*," that matches everything except a space. Well, this will also match anything because "*" means zero or more. The current version of sed can get unhappy with patterns like this, and generate errors like "Output line too long" or even run forever. I consider this a bug, and have reported this to Sun. As a work-around, you must avoid matching the null string when using the "g" flag to sed. A work-around example is: "[^ ][^ ]*." The following will put parenthesis around the first word:
  
-sed 's/[^ ]*/(&)/' <old >new+  sed 's/[^ ]*/(&)/' <old >new
  
 If you want it to make changes for every word, add a "g" after the last delimiter and use the work-around: If you want it to make changes for every word, add a "g" after the last delimiter and use the work-around:
  
-sed 's/[^ ][^ ]*/(&)/g' <old >new +  sed 's/[^ ][^ ]*/(&)/g' <old >new 
-Is sed recursive?+ 
 +===== Is sed recursive? =====
  
 Sed only operates on patterns found in the in-coming data. That is, the input line is read, and when a pattern is matched, the modified output is generated, and the rest of the input line is scanned. The "s" command will not scan the newly created output. That is, you don't have to worry about expressions like: Sed only operates on patterns found in the in-coming data. That is, the input line is read, and when a pattern is matched, the modified output is generated, and the rest of the input line is scanned. The "s" command will not scan the newly created output. That is, you don't have to worry about expressions like:
Ligne 142: Ligne 143:
  
 This will not cause an infinite loop. If a second "s" command is executed, it could modify the results of a previous command. I will show you how to execute multiple commands later. This will not cause an infinite loop. If a second "s" command is executed, it could modify the results of a previous command. I will show you how to execute multiple commands later.
-/1, /2, etc. Specifying which occurrence+ 
 +===== /1, /2, etc. Specifying which occurrence =====
  
 With no flags, the first pattern is changed. With the "g" option, all patterns are changed. If you want to modify a particular pattern that is not the first one on the line, you could use "\(" and "\)" to mark each pattern, and use "\1" to put the first pattern back unchanged. This next example keeps the first word on the line but deletes the second: With no flags, the first pattern is changed. With the "g" option, all patterns are changed. If you want to modify a particular pattern that is not the first one on the line, you could use "\(" and "\)" to mark each pattern, and use "\1" to put the first pattern back unchanged. This next example keeps the first word on the line but deletes the second:
  
-sed 's/\([a-zA-Z]*\) \([a-zA-Z]*\) /\1 /' <old >new+  sed 's/\([a-zA-Z]*\) \([a-zA-Z]*\) /\1 /' <old >new
  
 Yuck. There is an easier way to do this. You can add a number after the substitution command to indicate you only want to match that particular pattern. Example: Yuck. There is an easier way to do this. You can add a number after the substitution command to indicate you only want to match that particular pattern. Example:
  
-sed 's/[a-zA-Z]* //2' <old >new+  sed 's/[a-zA-Z]* //2' <old >new
  
 You can combine a number with the g (global) flag. For instance, if you want to leave the first world alone alone, but change the second, third, etc. to DELETED, use /2g: You can combine a number with the g (global) flag. For instance, if you want to leave the first world alone alone, but change the second, third, etc. to DELETED, use /2g:
  
-sed 's/[a-zA-Z]* /DELETED /2g' <old >new+  sed 's/[a-zA-Z]* /DELETED /2g' <old >new
  
 Don't get /2 and \2 confused. The /2 is used at the end. \2 is used in inside the replacement field. Don't get /2 and \2 confused. The /2 is used at the end. \2 is used in inside the replacement field.
Ligne 160: Ligne 162:
 Note the space after the "*" character. Without the space, sed will run a long, long time. (Note: this bug is probably fixed by now.) This is because the number flag and the "g" flag have the same bug. You should also be able to use the pattern Note the space after the "*" character. Without the space, sed will run a long, long time. (Note: this bug is probably fixed by now.) This is because the number flag and the "g" flag have the same bug. You should also be able to use the pattern
  
-sed 's/[^ ]*//2' <old >new+  sed 's/[^ ]*//2' <old >new
  
 but this also eats CPU. If this works on your computer, and it does on some Unix systems, you could remove the encrypted password from the password file: but this also eats CPU. If this works on your computer, and it does on some Unix systems, you could remove the encrypted password from the password file:
  
-sed 's/[^:]*//2' </etc/passwd >/etc/password.new+  sed 's/[^:]*//2' </etc/passwd >/etc/password.new
  
 But this didn't work for me the time I wrote thise. Using "[^:][^:]*" as a work-around doesn't help because it won't match an non-existent password, and instead delete the third field, which is the user ID! Instead you have to use the ugly parenthesis: But this didn't work for me the time I wrote thise. Using "[^:][^:]*" as a work-around doesn't help because it won't match an non-existent password, and instead delete the third field, which is the user ID! Instead you have to use the ugly parenthesis:
  
-sed 's/^\([^:]*\):[^:]:/\1::/' </etc/passwd >/etc/password.new+  sed 's/^\([^:]*\):[^:]:/\1::/' </etc/passwd >/etc/password.new
  
 You could also add a character to the first pattern so that it no longer matches the null pattern: You could also add a character to the first pattern so that it no longer matches the null pattern:
  
-sed 's/[^:]*:/:/2' </etc/passwd >/etc/password.new+  sed 's/[^:]*:/:/2' </etc/passwd >/etc/password.new
  
 The number flag is not restricted to a single digit. It can be any number from 1 to 512. If you wanted to add a colon after the 80th character in each line, you could type: The number flag is not restricted to a single digit. It can be any number from 1 to 512. If you wanted to add a colon after the 80th character in each line, you could type:
  
-sed 's/./&:/80' <file >new+  sed 's/./&:/80' <file >new
  
 You can also do it the hard way by using 80 dots: You can also do it the hard way by using 80 dots:
  
-sed 's/^................................................................................/&:/' <file >new +  sed 's/^................................................................................/&:/' <file >new 
-/p - print+ 
 +===== /p - print =====
  
 By default, sed prints every line. If it makes a substitution, the new text is printed instead of the old one. If you use an optional argument to sed, "sed -n," it will not, by default, print any new lines. I'll cover this and other options later. When the "-n" option is used, the "p" flag will cause the modified line to be printed. Here is one way to duplicate the function of grep with sed: By default, sed prints every line. If it makes a substitution, the new text is printed instead of the old one. If you use an optional argument to sed, "sed -n," it will not, by default, print any new lines. I'll cover this and other options later. When the "-n" option is used, the "p" flag will cause the modified line to be printed. Here is one way to duplicate the function of grep with sed:
  
-sed -n 's/pattern/&/p' <file +  sed -n 's/pattern/&/p' <file 
-Write to a file with /w filename+   
 +===== Write to a file with /w filename =====
  
 There is one more flag that can follow the third delimiter. With it, you can specify a file that will receive the modified data. An example is the following, which will write all lines that start with an even number to the file even: There is one more flag that can follow the third delimiter. With it, you can specify a file that will receive the modified data. An example is the following, which will write all lines that start with an even number to the file even:
  
-sed -n 's/^[0-9]*[02468] /&/w even' <file+  sed -n 's/^[0-9]*[02468] /&/w even' <file
  
 In this example, the output file isn't needed, as the input was not modified. You must have exactly one space between the w and the filename. You can also have ten files open with one instance of sed. This allows you to split up a stream of data into separate files. Using the previous example combined with multiple substitution commands described later, you could split a file into ten pieces depending on the last digit of the first number. You could also use this method to log error or debugging information to a special file. In this example, the output file isn't needed, as the input was not modified. You must have exactly one space between the w and the filename. You can also have ten files open with one instance of sed. This allows you to split up a stream of data into separate files. Using the previous example combined with multiple substitution commands described later, you could split a file into ten pieces depending on the last digit of the first number. You could also use this method to log error or debugging information to a special file.
-Combining substitution flags+ 
 +===== Combining substitution flags =====
  
 You can combine flags when it makes sense. Also "w" has to be the last flag. For example the following command works: You can combine flags when it makes sense. Also "w" has to be the last flag. For example the following command works:
  
-sed -n 's/a/A/2pw /tmp/file' <old >new+  sed -n 's/a/A/2pw /tmp/file' <old >new
  
 Next I will discuss the options to sed, and different ways to invoke sed. Next I will discuss the options to sed, and different ways to invoke sed.
-Arguments and invocation of sed+ 
 +===== Arguments and invocation of sed =====
  
 previously, I have only used one substitute command. If you need to make two changes, and you didn't want to read the manual, you could pipe together multiple sed commands: previously, I have only used one substitute command. If you need to make two changes, and you didn't want to read the manual, you could pipe together multiple sed commands:
  
-sed 's/BEGIN/begin/' <old | sed 's/END/end/' >new+  sed 's/BEGIN/begin/' <old | sed 's/END/end/' >new
  
 This used two processes instead of one. A sed guru never uses two processes when one can do. This used two processes instead of one. A sed guru never uses two processes when one can do.
-Multiple commands with -e command+ 
 +===== Multiple commands with -e command =====
  
 One method of combining multiple commands is to use a -e before each command: One method of combining multiple commands is to use a -e before each command:
  
-sed -e 's/a/A/' -e 's/b/B/' <old >new+  sed -e 's/a/A/' -e 's/b/B/' <old >new
  
 A "-e" isn't needed in the earlier examples because sed knows that there must always be one command. If you give sed one argument, it must be a command, and sed will edit the data read from standard input. A "-e" isn't needed in the earlier examples because sed knows that there must always be one command. If you give sed one argument, it must be a command, and sed will edit the data read from standard input.
  
 Also see Quoting multiple sed lines in the Bourne shell Also see Quoting multiple sed lines in the Bourne shell
-Filenames on the command line+ 
 +===== Filenames on the command line =====
  
 You can specify files on the command line if you wish. If there is more than one argument to sed that does not start with an option, it must be a filename. This next example will count the number of lines in three files that don't begin with a "#:" You can specify files on the command line if you wish. If there is more than one argument to sed that does not start with an option, it must be a filename. This next example will count the number of lines in three files that don't begin with a "#:"
francois/an_introduction_to_sed.1368520254.txt.gz · Dernière modification : (modification externe)

Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki