You read bytes as UTF-8 and they aren't. Either the file genuinely isn't UTF-8 (Windows-1252/Latin-1/UTF-16), or it is but a split read broke a multi-byte character. Decode explicitly, detect when unsure, and never read with errors ignored on data that matters.
Default open() assumes UTF-8. Excel exports (cp1252), Windows text (BOM + utf-16), old DB dumps: byte 0xNN positions tell the family. file <name> and chardetect identify; the 0x92/0x93 range is the cp1252 smart-quote giveaway.
Reading fixed-size chunks can split a multi-byte UTF-8 sequence: the file IS valid, the chunk boundary isn't. Incremental decoders (codecs.getincrementaldecoder) handle chunk boundaries properly.
file -i data.csv ; python3 -c 'import chardet; print(chardet.detect(open("data.csv","rb").read(100000)))'
open('data.csv', encoding='cp1252') # or 'utf-16', 'latin-1' — explicit beats guessed
open('out.csv', 'w', encoding='utf-8') # the one-liner that prevents the next person's error
# codecs.getincrementaldecoder('utf-8')() per chunk — or read lines (never fixed byte counts) from text files
errors='replace' / errors='ignore' turn decode crashes into silent data corruption — acceptable for logs you're grepping, never for data you process or store. A BOM (utf-8-sig) causes a different invisible bug: the first column name gains \ufeff. If pandas/groupby mysteriously fails on the FIRST key only, that's the BOM.
Only the ones containing bytes outside UTF-8's rules (smart quotes, accented letters in cp1252). ASCII-only files decode fine as anything — the error appears exactly when non-ASCII data arrives.
For throwaway exploration, maybe; for anything persisted, no: it silently drops the problem bytes (and surrounding meaning). Decode with the true encoding — chardet on a sample is a 5-second answer.
Our most-documented failures, packaged as ready-to-ship starter kits: Docker, Kubernetes, and Terraform.
Browse the template store →One-time. Yours to modify. Instant download from the NinjaOps template store.