Data Files: Difference between revisions
Jump to navigation
Jump to search
No edit summary |
m →Example |
||
| (7 intermediate revisions by 3 users not shown) | |||
| Line 8: | Line 8: | ||
===== Lines ===== | ===== Lines ===== | ||
* Lines are separated by the '''\r\n''' line delimiter for better compatibility between operating systems. | * Lines are separated by the '''\r\n''' line delimiter for better compatibility between operating systems. | ||
* The line delimiter should also be added after the last line. This simplifies stream reading since all records (lines) are terminated. | * The line delimiter should also be added after the last line. This simplifies stream reading since all records (lines) are terminated. This allows for the use of a readline() function for acquiring a line. | ||
* The first line contains a header with column/field names. | * The first line contains a header with column/field names. | ||
===== Fields ===== | ===== Fields ===== | ||
* | * Fields are separated by the '''tab''' field delimiter, because they rarely occur in texts. This allows for the use of comma's and semicolons in sentences without using an escape character. | ||
* The field delimiter should | * The field delimiter should '''not''' be added after each line's last field. This allows for the use of a split() function for parsing a line. | ||
* The last field in a line must not be empty, because | * The last field in a line must not be empty, because it will show to parsers that the previous rule was obeyed. | ||
* Fields are | * Fields are never surrounded by a quoting character. | ||
* White space | * White space before or after field delimiters are considered part of a field. | ||
* There is no defined escape character. If your data can contain tabs, use a different field delimiter or file format. | * There is no defined escape character. If your data can contain tabs or newlines, use a different field delimiter or file format. | ||
===== Data ===== | ===== Data ===== | ||
| Line 32: | Line 32: | ||
</pre> | </pre> | ||
An example file can be downloaded here | An example file can be downloaded here {{Download link|media=Example.zip}} (sorry, it is zipped). | ||
== Parsing == | == Parsing == | ||
| Line 71: | Line 71: | ||
{| class="wikitable" | {| class="wikitable" | ||
|- | |- | ||
| | | File Extension || '''tsv''', csv, '''dat''', txt | ||
|- | |- | ||
| | | File Encoding || '''ASCII''', '''UTF-8''', UTF-16BE, UTF-16LE, UCS-4/UTF-32 | ||
|- | |- | ||
| | | [[wikipedia:Magic_number_(programming)|Magic Number]] || '''None''', [[wikipedia:Byte_order_mark|BOM]] | ||
|- | |- | ||
| | | Line Delimiter || \n, \r, '''\r\n''' | ||
|- | |- | ||
| | | Line Delimiter after Last Line || '''Yes''', No | ||
|- | |- | ||
| | | Field Delimiter || '''<tab>''', <comma> , <semicolon> | ||
|- | |- | ||
| | | Field Delimiter after Last Field || Yes, '''No''' | ||
|- | |- | ||
| | | Quoting Character || '''None''', ', " | ||
|- | |- | ||
| | | Escape QC by doubling || Yes, No | ||
|- | |- | ||
| | | Escape Character || '''None''', \ | ||
|- | |- | ||
| | | First Line Contains: || '''Header''', Data | ||
|- | |- | ||
| | | Empty Last Field in Line || Allowed, '''Not Allowed''' | ||
|- | |- | ||
| | | Whitespace Following Delimiter || '''Part of Field''', Excluded | ||
|- | |- | ||
| | | Decimal Separator || '''<dot>''', <comma> | ||
|- | |- | ||
| | | Thousands Separator || '''None''', <dot>, <space>, U+2009 | ||
|} | |} | ||
Note that tab characters and newlines cannot be present in field content. | Note that tab characters and newlines cannot be present in field content. | ||