fix: 兼容GBK不间断空格

This commit is contained in:
2026-08-07 16:34:18 +08:00
parent 8d681118ea
commit 8a3c273a04
5 changed files with 15 additions and 4 deletions
+1
View File
@@ -104,3 +104,4 @@
- All generated CSV files now use GBK without a UTF-8 BOM: the seven full source conversions, the merged summary, repaired history exports, and uploaded history files. `manifest.json`, API JSON, and database character sets remain unchanged.
- The encoding is centralized in `CSV_ENCODING`; tests verify GBK Chinese bytes, absence of the UTF-8 BOM, summary parsing, and history parsing.
- Production data contains non-breaking spaces (`U+00A0`), which Python's GBK codec cannot encode. CSV-only normalization converts them to ordinary spaces while keeping strict encoding for every other unsupported character, so unexpected data still fails visibly instead of losing text silently.