Fixing a bug with byte order marks

(alexwlchan.net)

5 points | by surprisetalk 5 days ago

4 comments

  • orangepanda 0 minutes ago
    > A byte order mark is a special use of the zero width no-break space character U+FEFF at the beginning of a text file

    Isnt BOM allowed to appear anywhere in the file, because of file concatenation?

  • flohofwoe 55 minutes ago
    Ugh, why are BOMs even still a thing in the 21st century? E.g. when will Windows finally arrive in the late 1990s and switch to UTF-8 for everything?

    The UTF-8 BOM is especially bizarre because UTF-8 is completely endian-agnostic (so this "UTF-8 Byte Order Mark" is at most an indicator that this file is UTF-8 encoded, but guess what? Outside the Windows bubble, all text files are UTF-8 anyway).

  • mr_mitm 50 minutes ago
    I hate the BOM so much. It causes so many subtle issues and I don't even understand why it's needed.
  • remnavi 53 minutes ago
    The csv.DictReader version of this bug is the one that got me. Open a UTF-8-with-BOM CSV as plain utf-8 and your first column name comes back with an invisible U+FEFF glued to the front of it, so every lookup on that one field fails while the other columns behave perfectly. You end up suspecting the source data, the header row, anything except the encoding.

    json.loads has a similar tell: a leading BOM gives you "Expecting value: line 1 column 1 (char 0)", which reads like malformed JSON when the JSON is fine.

    Same fix, utf-8-sig at the open, and I think the lesson generalises past subtitles: deal with encoding once at the I/O boundary so nothing downstream ever needs to know a BOM existed. Mid-file BOMs like yours are what it looks like when that leaks.