Line 8: |
Line 8: |
| ===== Lines ===== | | ===== Lines ===== |
| * Lines are separated by the '''\r\n''' line delimiter for better compatibility between operating systems. | | * Lines are separated by the '''\r\n''' line delimiter for better compatibility between operating systems. |
− | * The line delimiter should also be added after the last line. This simplifies stream reading since all records (lines) are terminated. | + | * The line delimiter should also be added after the last line. This simplifies stream reading since all records (lines) are terminated. This allows for the use of a readline() function for acquiring a line. |
| * The first line contains a header with column/field names. | | * The first line contains a header with column/field names. |
| | | |
| ===== Fields ===== | | ===== Fields ===== |
− | * Field are separated by the '''tab''' field delimiter, because they rarely occur in texts and therefore require no escaping. | + | * Fields are separated by the '''tab''' field delimiter, because they rarely occur in texts. This allows for the use of comma's and semicolons in sentences without using an escape character. |
| * The field delimiter should '''not''' be added after each line's last field. This allows for the use of a split() function for parsing a line. | | * The field delimiter should '''not''' be added after each line's last field. This allows for the use of a split() function for parsing a line. |
| * The last field in a line must not be empty, because it will show to parsers that the previous rule was obeyed. | | * The last field in a line must not be empty, because it will show to parsers that the previous rule was obeyed. |
| * Fields are never surrounded by a quoting character. | | * Fields are never surrounded by a quoting character. |
| * White space before or after field delimiters are considered part of a field. | | * White space before or after field delimiters are considered part of a field. |
− | * There is no defined escape character. If your data can contain tabs, use a different field delimiter or file format. | + | * There is no defined escape character. If your data can contain tabs or newlines, use a different field delimiter or file format. |
| | | |
| ===== Data ===== | | ===== Data ===== |
Line 71: |
Line 71: |
| {| class="wikitable" | | {| class="wikitable" |
| |- | | |- |
− | | file extension || '''tsv''' || csv || '''dat''' || txt | + | | File Extension || '''tsv''', csv, '''dat''', txt |
| |- | | |- |
− | | file extension || '''ascii''' || '''UTF-8''' || UTF-16BE || UTF-16LE || UCS-4/UTF-32 | + | | File Encoding || '''ASCII''', '''UTF-8''', UTF-16BE, UTF-16LE, UCS-4/UTF-32 |
| |- | | |- |
− | | magic number || '''None''' || <BOM> | + | | [[wikipedia:Magic_number_(programming)|Magic Number]] || '''None''', [[wikipedia:Byte_order_mark|BOM]] |
| |- | | |- |
− | | line delimiter || \n || \r || '''\r\n''' | + | | Line Delimiter || \n, \r, '''\r\n''' |
| |- | | |- |
− | | line delimiter after last line || no || '''yes''' | + | | Line Delimiter after Last Line || '''Yes''', No |
| |- | | |- |
− | | field delimiter || '''<tab>''' || , || ; | + | | Field Delimiter || '''<tab>''', <comma> , <semicolon> |
| |- | | |- |
− | | field delimiter after last field || '''no''' || yes | + | | Field Delimiter after Last Field || Yes, '''No''' |
| |- | | |- |
− | | quoting character || '''None''' || " || ' | + | | Quoting Character || '''None''', ', " |
| |- | | |- |
− | | escape qc by doubling || no || yes | + | | Escape QC by doubling || Yes, No |
| |- | | |- |
− | | escape character || '''none''' || \ | + | | Escape Character || '''None''', \ |
| |- | | |- |
− | | first line || '''contains header''' || contains data | + | | First Line Contains: || '''Header''', Data |
| |- | | |- |
− | | last field in line || '''must not be empty''' || may be empty | + | | Empty Last Field in Line || Allowed, '''Not Allowed''' |
| |- | | |- |
− | | whitespace following delimiter || '''part of field''' || not part of field | + | | Whitespace Following Delimiter || '''Part of Field''', Excluded |
| |- | | |- |
− | | decimal separator || '''.''' || , | + | | Decimal Separator || '''<dot>''', <comma> |
| |- | | |- |
− | | thousands separator || '''none''' || . || ␣ || U+2009 | + | | Thousands Separator || '''None''', <dot>, <space>, U+2009 |
| |} | | |} |
| Note that tab characters and newlines cannot be present in field content. | | Note that tab characters and newlines cannot be present in field content. |