Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
You've heard of CSV files, but have you heard of CCSV files? (robida.net)
35 points by surprisetalk 22 hours ago | hide | past | favorite | 28 comments
 help



This comes up every so often. I think there are good reasons it hasn't caught on. Here's a relevant comment from ten years ago:

> whole point is to be roughly human-readable and using non-printing characters defeats that. You can't even easily enter these things via the command line.

> If we're abandoning human-readability, why even bother with ASCII? Just use a binary format. Has anyone actually used ASCII unit and record separator delimiters successfully? I'd be curious about what advantages they had over a binary format, even just a protobuf or Thrift serialized form. If we want to preserve schemalessness, there's stuff like Sereal.

--- arjie, June 8, 2016

https://news.ycombinator.com/item?id=11862769

Here's another one from more than twelve years ago:

> I've done this.

> Everybody hated it. Most text editors don't display anything useful with these characters (either hiding them altogether or showing a useless "uknown" placeholder), and spreadhseet tools don't support the record separator (although they all let you provide a custom entry separator so the "unit" separator can work). Besides the obvious problem that there's no easy way to type the darned things when somebody hand-edits the file.

--- Pxtl, March 26, 2014

https://news.ycombinator.com/item?id=7474600


I think there is a deeper question: "record separator" and "unit separator" chars are a good idea, but why are they invisible? I agree that in the current form, they are mostly unusable.

What was the idea of the designers how they should be used?


ASCII was created in the 1960s. That's around the time when the industry wasn't really sure if all files were just going to be a stream of bytes that were completely up to programs to interpret, or if it was the OS's responsibility to enforce a database-like structure on all disk-like I/O.

Yahoo access logs used control characters to separate fields beyond a fixed width field, but it wasn't the record separators. It used ctrl-E to separate fields, which had a one character identifier. Some of the fields had sub-fields which were separated ctrl-F.

Slide 22 https://www.radwin.org/michael/talks/yapache-oscon2006.pdf


On displaying control characters, as of now, at least VSCode and Sublime Text do show them clearly. VSCode uses Unicode control pictures “␄” with a red background, while Sublime shows plain-text “<EOT>” in grey. You can also copy-paste - the real control character lands on your clipboard. Otherwise the rest of it stands, especially for a non-developer using Notepad or TextEdit.

This is trading off ongoing usability for one-time developer convenience. The example given in the post would also struggle with a large file as it loads the entire contents into memory, while having Python feed you lines allows it to read in chunks. Plenty of accurate, unit-tested CSV parsers exist in every language, it’s fine to use one and be done with it. Tabular data formats are a solved problem.


The problem in practice with CCSV files, is they are no longer human editable, nor will tools like Excel automatically parsed them (maybe, I haven't checked in ~10 years, however I doubt CSV parsing has improved at all).

If the files aren't human editable anymore, why stick with CSV at all?


I think similar way - CSV should be abandoned and we should use different, better format.

I mean no one is arguing that it is the best or even an OK format for serious data storage. Other solutions exist for a reason. But it’s evergreen and won’t go away precisely because of its simplicity. You will be able to use a CSV more or less identically in 2050, no questions asked.

It’s inherently universal because it’s text-based, and I think you underrate the occasional need to inspect it yourself. And when you do, it’s nicer to view than JSON to throw out an example (arguably JSON and variants need their own bespoke file viewer!)

It’s like saying markdown should be abandoned because text editors do all the formatting for you anyways and different markdown parsers exist and don’t always play nice. Or like saying txt should be abandoned in favor of rtf, even. Okay, sure, not invalid points. But in practice? Simplicity is often a virtue.


Firefox displays .json files decently if you ever need something in a pinch.

excel can edit csvs, definitely, no problem at all. It will lose any formatting and column size adjustments if you save back as csv, but that's expected.

This is generally called ASV for ASCII-separated values, or USV for Unicode-separated values (e.g. U+241E, which is different from the ASCII U+001E equivalent).

Also, as wonderful as this would be, as other comments say, no one uses it because no one uses it.


I think we hugged it to death.

https://archive.is/4anlI


The article seems to be unaware that Python has a built-in csv module.

I just use tsv. No quoted fields required because none of my data ever has tabs

Why was a comma originally chosen over, I dunno, say a pipe character?

Yeah, this is something frustrating to me. A comma is just way too common of a symbol used in everyday text. So if you're using something other than numerical data like strings. Quotes are another thing where trying to escape quotes escaping commas decreases the human readability. CSV parsers can usually designate the delimiter and escape characters, so you can use a more unique delimiter. Will the recipient of your file parse correctly or just assume it's a single column file?

Because English. If you're writing a data file and one row of data elements is like "one, red, 3.14, Tuesday," why wouldn't you write exactly that in your ASCII text file? And when you do, that's CSV!

If you've dealt with the headache that can be FIX, you'll realize the annoyance this can be because unless your tooling is very aware you're doing this (it probably isn't, unless you've built your own) everything looks like garbage until you tune everything to make it pretty.

Works internally, breaks externally, which half of the benefit of a CSV is that you can give just about anyone the file and they can work with it.


PICK D3 has been doing similar since... well, a very long time. It uses chr(254), chr(253) and chr(252) for the varying depths of records. I always rather enjoyed working with it in the 2000s. Elegant in its simplicity, functionally practical (for the workloads we had at the time at least)

The Canadian government uses tab-separated values quite a lot.

These control characters do sometimes appear in CSV files in the wild. The suggested method of parsing is unsafe and lazy.

We deal with mostly business data, doing migrations. | often works out to be a nice, non-conflicting delimiter.

One can pipe it to an expression for viewing purposes:

    sed $'s/\x1f/,/g; s/\x1e/\\\n/g; s/\x04//g'

    awk 'BEGIN { RS="\036"; FS="\037"; OFS="," } { sub(/\004$/, ""); $1=$1; print }'

I thought this was going to be about comma-separated CSV files.

Then that's just CSV?

It's turtles all the way down, mate.

Turtle-separated values? I thought TSV meant tabs.

Seems to be down



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: