<- Back
Comments (81)
- devyGraydon Hoare, the creator of the Rust programming language, wrote his seminal piece on text in 2014. In it, he said "text is the most powerful, useful, effective communication technology ever, period."[1] Text is durable.[1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)[2] https://hn.algolia.com/?q=always+bet+on+text[3] https://news.ycombinator.com/item?id=26164001[4] https://news.ycombinator.com/item?id=8451271[5] https://news.ycombinator.com/item?id=10284202[6] https://news.ycombinator.com/item?id=12815829
- boomlindeThe article briefly addresses the problem, but it's pretty fun how different "plain text" looks throughout history and in different domains.For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.
- drhagenThis is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0].[0] Windows newlines not withstanding.
- wpollockI too love plain text data, but absent a universally accepted standard, meta data is needed to avoid guessing a document's charset, encoding, and line ending convention.Associating such meta data with text files wasn't simple in the beginning, but that changed in the 1990s. Nearly all filesystems since then allow meta data: POSIX compliant filesystems support extended attributes, mostly used for security meta data. Even ancient HFS (Apple Mac) supports meta data. NTFS supports it with alternative file streams.While possible, no system I know of attempts to use any such meta data, probably due to a lack of standards. A pity.
- kerblangWhat, no mention of our old friend, ASCII-armored Base64? For shame! Everything can be text with Base64, including things that have absolutely no business being text! Best of all, in light of popular widespread abuse of every available resource, Base64 is comparatively efficient! Bring on the petabytes! Yay and I'm not being completely sarcastic
- andsoitisThe more easily you can access meaning from symbols without intermediate tools, the more durable and resilient.Symbols on paper is best.Plain text in a file that can be opened by any computer comes close, but needs a tool.Fancy file formats that need not only hardware but also special software are the worst.However, there's an important tradeoff, which is the fancier formats can present information in ways mere symbols might struggle with, and can also use interactivity to improve understanding.
- spcebarFun timing on this plaintext conversation. I built a little browser based plaintext playwriting app this weekend for a little weekend project. Found myself really hating 1. how clunk screenwriting software can be and 2. How unsharable the files are.With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.
- shmolyneauxI love plain text, but the calculus of being able to access the content in 50 years is much more interesting in the age of agents. They can infer the meaning of structured data without schemas, decompressed archives, find embedded files, etc. It's not perfect, but the durability of binary formats is better now than it has ever been.It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).
- frollogastonI greatly prefer plaintext in most cases where Markdown ends up being used. WYSIWYG is more useful than formatting in these cases. Most of the formatting makes it harder to read anyway. Some say you can ignore MD and treat it like plaintext... not so when you use a newline.Oh and now you could have LLMs write the Markdown, but chances are no human reads that, or even if you do, the formatting is going to be insane. I have to keep telling Claude to give me a .txt instead of .md when asking for a context dump. Maybe .md is popular for agent readmes because they optimize around its headings.
- njarboeI work on scientific data repositories [1] and pushed for plain text files as our archiving file format two decades ago as the best long term format. We implemented it on our data repositories. Very human readable also. The datasets are small and table based, so this works well. We latter started archiving 2-D image data and that creates very large files. We are still looking for an elegant solutions for those datasets.[1] https://earthref.org/FIESTA/
- CrimsonCapeHijacking the plain text discussion to ask what the best tools are to convert plain text to lexed/parsed output?I'm assuming any tool in this regard would expect the user to write an EBNF grammar.I found ANTLR to be nice but it's way too convoluted to use as a tool with Java dependencies and seems to be stagnating. And tree sitter just is too convoluted and requires the added C/C++ overhead to understand how to use it.
- anvuongASCII was amazingly efficient for conveying English text. Then we needed to encode multi languages and emojis, the resulting Unicode is just a mess.
- pratikdeoghareI like plain text. I tolerate markdown. Markdown was designed as a shorthand for html. It gets stretched to do other stuff.I created Brashtag [1]. It is simpler than markdown.[1] https://github.com/PratikDeoghare/brashtag
- mythrilkey30I dropped the first 3 paragraphs into pangram and it is 100% AI generated. Not reading and moving on.
- GrimetonASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.Unicode are multibyte characters with variable byte length and endianess at play. If you read it wrong or guess the length wrong your results might be anything but useful.
- vatsachakProbably because plain text can encode any form of distilled data. Technically our DNA can be plain text lol
- ollienTangential to the actual point of the post, but the talk "Plain text? Really?" by Dylan Beattie[1] is one of my favorite talks. It does a great job capturing the problems with something that "seems" so simple. I think the author of the blog is well aware of these, though :)[1] https://www.youtube.com/watch?v=_mZBa3sqTrI
- dimiprasakisPrediction: RSS is coming back stronger than ever
- olexsmirI feel obligated to mention ledger[0] and hledger[1], those are plain text accounting software, well, as you can guess from their names, they allow you to do personal accounting in plain text.0: https://ledger-cli.org1: https://hledger.org
- leftnodeBecause no single person or entity can own it.
- m463I think text is getting a big boost from the age of ai.
- jaekwonthat's why gno.land contracts render to markdown as the standard. you can browse the world through your terminal.imagine the world wide web but markdown (and decentralized). gno.land is that.
- ummonk> The Unicode Standard defines plain text essentially as a sequence of character codes, without the additional formatting information associated with rich text. Fonts, colors, layout, and similar presentation details belong somewhere else.Nope. Skintones are part of Unicode. It also has 33 control characters from ASCII including one that rings a bell... There are also numerous characters added by Unicode that are literally called "layout controls".If you want a format that provides purely semantic information, then "plaintext" doesn't fit the bill.
- charcircuit>There are not many computer file formats I would trust to still be readable fifty years from now, but plain text is one of them.Just the other day I had opus read a 10 year old proprietary file format. The idea that we will lose the ability to use file formats is not consistent with reality.
- mythrilkey30I pasted the first 3 paragraphs of this into pangram and it's 100% AI generated. Not reading, moving on. We really shouldn't be promoting slop on hacker news
- Koshkin"A picture is worth a thousand words"
- FLeXMurphy@dang @tomhowMy understanding is that HN has started incorporating AI tooling in the back-end to scan for LLM-generated submissions and source content, in an effort to discourage it and encourage human-made content (with human discussions, one would hope). Why do we keep having these kinds of articles every day? Entire "apps" are generated - see the lighthouse one - and submitted, and are plain-as-day LLM-generated nonsense.Anyway, flagged.