fix: 仅对储存历史CSV使用GBK

This commit is contained in:
2026-08-07 16:37:20 +08:00
parent 8a3c273a04
commit 8e96b3a24b
5 changed files with 36 additions and 19 deletions
+1 -1
View File
@@ -48,7 +48,7 @@ After the database succeeds, the same eleven result columns are uploaded as GBK
The filename time comes from normalized `metric_time`, never the server clock. If the database already contains the target time but its history CSV is missing or empty, the script exports that time from the database and repairs the history file without reprocessing source data. The filename time comes from normalized `metric_time`, never the server clock. If the database already contains the target time but its history CSV is missing or empty, the script exports that time from the database and repairs the history file without reprocessing source data.
All generated CSV files use GBK without a UTF-8 BOM, including the seven converted source files, the merged summary, and the uploaded history file. Source non-breaking spaces (`U+00A0`) are normalized to ordinary spaces because GBK cannot encode them; other unsupported characters still fail the run instead of being silently replaced. `manifest.json` remains UTF-8 JSON. Only the final history CSV uploaded to Metrix Storage uses GBK without a UTF-8 BOM. Source non-breaking spaces (`U+00A0`) are normalized to ordinary spaces in this history file because GBK cannot encode them; other unsupported characters still fail visibly. The seven local converted CSV files and local merged summary remain UTF-8 with BOM, and `manifest.json` remains UTF-8 JSON.
CGI is generated with fixed rules: CGI is generated with fixed rules:
+4 -4
View File
@@ -100,8 +100,8 @@
- Offline package SHA-256 is `74EEA84058C5F14E7108B840561C070DC47752D67FEC68262F0CC6C78EE8EC0D`. SSH deployment run `984ffc6f98024a2d81b63d4359aabc71` processed `2026-08-07 07:00:00` with 820 rows and deleted 1,137 rows from the prior database time. - Offline package SHA-256 is `74EEA84058C5F14E7108B840561C070DC47752D67FEC68262F0CC6C78EE8EC0D`. SSH deployment run `984ffc6f98024a2d81b63d4359aabc71` processed `2026-08-07 07:00:00` with 820 rows and deleted 1,137 rows from the prior database time.
- Production verification matched 616 current CGI values to the previous period and left both previous fields NULL for the remaining 204 CGI values. `prev_nearby_count` matched exactly; `prev_interference_dbm` uses the table's existing `DECIMAL(10,3)` precision. Repeat run `6b38030f6d9e4ab1a0ee6a4ba0dc5d6e` returned `status=skipped`. - Production verification matched 616 current CGI values to the previous period and left both previous fields NULL for the remaining 204 CGI values. `prev_nearby_count` matched exactly; `prev_interference_dbm` uses the table's existing `DECIMAL(10,3)` precision. Repeat run `6b38030f6d9e4ab1a0ee6a4ba0dc5d6e` returned `status=skipped`.
## 2026-08-07: GBK CSV output ## 2026-08-07: GBK Storage history output
- All generated CSV files now use GBK without a UTF-8 BOM: the seven full source conversions, the merged summary, repaired history exports, and uploaded history files. `manifest.json`, API JSON, and database character sets remain unchanged. - Only final history CSV files uploaded to Metrix Storage use GBK without a UTF-8 BOM, including database-based history repairs. The seven local source conversions and local merged summary remain UTF-8 with BOM; `manifest.json`, API JSON, and database character sets remain unchanged.
- The encoding is centralized in `CSV_ENCODING`; tests verify GBK Chinese bytes, absence of the UTF-8 BOM, summary parsing, and history parsing. - Production data contains non-breaking spaces (`U+00A0`), which Python's GBK codec cannot encode. History-output normalization converts them to ordinary spaces while keeping strict encoding for every other unsupported character.
- Production data contains non-breaking spaces (`U+00A0`), which Python's GBK codec cannot encode. CSV-only normalization converts them to ordinary spaces while keeping strict encoding for every other unsupported character, so unexpected data still fails visibly instead of losing text silently. - Encoding behavior is separated through `LOCAL_CSV_ENCODING` and `HISTORY_CSV_ENCODING`; tests verify the local UTF-8 BOM and GBK history parsing independently.
+1 -1
View File
@@ -8,7 +8,7 @@
- Only `Sheet0` is converted and merged. The `指标(计数器)` sheet is metadata and is not included in the summary. - Only `Sheet0` is converted and merged. The `指标(计数器)` sheet is metadata and is not included in the summary.
- NR CGI is derived as `{gNBplmn}-{gNBId}-{cellId}`. - NR CGI is derived as `{gNBplmn}-{gNBId}-{cellId}`.
- 4G CGI is derived as `460-00-{eNodeBID}-{小区ID}`, using the equivalent node and cell column names in each source schema. - 4G CGI is derived as `460-00-{eNodeBID}-{小区ID}`, using the equivalent node and cell column names in each source schema.
- All output CSV files use GBK without a UTF-8 BOM. Non-breaking spaces are normalized to ordinary spaces; other unsupported characters remain encoding errors. `manifest.json` remains UTF-8 JSON. - Only history CSV files uploaded to Metrix Storage use GBK without a UTF-8 BOM. Local converted and summary CSV files remain UTF-8 with BOM; `manifest.json` remains UTF-8 JSON.
- Source files are read-only. The script writes only below its configured output directory. - Source files are read-only. The script writes only below its configured output directory.
- MySQL is the scheduled-run output. The table uses one `metric_time DATETIME` column containing the source KPI start time. - MySQL is the scheduled-run output. The table uses one `metric_time DATETIME` column containing the source KPI start time.
- A successful run transactionally refreshes the selected hour and deletes rows for all other hours. A selected hour older than the newest database hour is rejected to prevent data rollback. - A successful run transactionally refreshes the selected hour and deletes rows for all other hours. A selected hour older than the newest database hour is rejected to prevent data rollback.
+21 -7
View File
@@ -82,7 +82,8 @@ LTE_PLMN = "460-00"
DATABASE_NAME = "interference_etl" DATABASE_NAME = "interference_etl"
DATABASE_TABLE = "interference_hourly_summary" DATABASE_TABLE = "interference_hourly_summary"
DEFAULT_HISTORY_ROOT = f"{DEFAULT_ROOT}/干扰历史数据" DEFAULT_HISTORY_ROOT = f"{DEFAULT_ROOT}/干扰历史数据"
CSV_ENCODING = "gbk" LOCAL_CSV_ENCODING = "utf-8-sig"
HISTORY_CSV_ENCODING = "gbk"
HEADER_FDD = ( HEADER_FDD = (
"开始时间", "开始时间",
@@ -806,7 +807,10 @@ def process(
stored_rows = store.rows_for_time(metric_time) stored_rows = store.rows_for_time(metric_time)
if not stored_rows: if not stored_rows:
raise ProcessingError(f"Database contains no rows for {metric_time_text}") raise ProcessingError(f"Database contains no rows for {metric_time_text}")
history_path = history.upload(metric_time, dict_csv_bytes(SUMMARY_HEADER, stored_rows)) history_path = history.upload(
metric_time,
dict_csv_bytes(SUMMARY_HEADER, stored_rows, HISTORY_CSV_ENCODING),
)
print("status=history_repaired") print("status=history_repaired")
print(f"target_time={metric_time_text}") print(f"target_time={metric_time_text}")
print(f"history_output={history_path}") print(f"history_output={history_path}")
@@ -867,7 +871,10 @@ def process(
if store is not None: if store is not None:
database_result = store.replace_latest(summary_rows) database_result = store.replace_latest(summary_rows)
if history is not None: if history is not None:
history_path = history.upload(metric_time, dict_csv_bytes(SUMMARY_HEADER, summary_rows)) history_path = history.upload(
metric_time,
dict_csv_bytes(SUMMARY_HEADER, summary_rows, HISTORY_CSV_ENCODING),
)
history_result = {"enabled": True, "path": history_path} history_result = {"enabled": True, "path": history_path}
matched_coordinates = sum(bool(row["longitude"] and row["latitude"]) for row in summary_rows) matched_coordinates = sum(bool(row["longitude"] and row["latitude"]) for row in summary_rows)
manifest = { manifest = {
@@ -1087,24 +1094,31 @@ def sql_decimal_literal(value: str, field: str, nullable: bool = False) -> str:
def write_csv(path: Path, header: tuple[str, ...], rows: list[tuple[object, ...]]) -> None: def write_csv(path: Path, header: tuple[str, ...], rows: list[tuple[object, ...]]) -> None:
with path.open("w", encoding=CSV_ENCODING, newline="") as file: with path.open("w", encoding=LOCAL_CSV_ENCODING, newline="") as file:
writer = csv.writer(file) writer = csv.writer(file)
writer.writerow(header) writer.writerow(header)
for row in rows: for row in rows:
writer.writerow(normalize_csv_cell(value) for value in row) writer.writerow(normalize_cell(value) for value in row)
def write_dict_csv(path: Path, header: tuple[str, ...], rows: list[dict[str, str]]) -> None: def write_dict_csv(path: Path, header: tuple[str, ...], rows: list[dict[str, str]]) -> None:
path.write_bytes(dict_csv_bytes(header, rows)) path.write_bytes(dict_csv_bytes(header, rows))
def dict_csv_bytes(header: tuple[str, ...], rows: list[dict[str, str]]) -> bytes: def dict_csv_bytes(
header: tuple[str, ...],
rows: list[dict[str, str]],
encoding: str = LOCAL_CSV_ENCODING,
) -> bytes:
output = io.StringIO(newline="") output = io.StringIO(newline="")
writer = csv.DictWriter(output, fieldnames=header, extrasaction="raise") writer = csv.DictWriter(output, fieldnames=header, extrasaction="raise")
writer.writeheader() writer.writeheader()
for row in rows: for row in rows:
if encoding == HISTORY_CSV_ENCODING:
writer.writerow({column: normalize_csv_cell(row[column]) for column in header}) writer.writerow({column: normalize_csv_cell(row[column]) for column in header})
return output.getvalue().encode(CSV_ENCODING) else:
writer.writerow(row)
return output.getvalue().encode(encoding)
def ensure_scoped(root: Path, target: Path) -> None: def ensure_scoped(root: Path, target: Path) -> None:
+8 -5
View File
@@ -40,9 +40,8 @@ class PipelineTest(unittest.TestCase):
self.assertEqual(len(converted), 7) self.assertEqual(len(converted), 7)
summary_path = result / f"interference_summary_{TARGET_KEY}.csv" summary_path = result / f"interference_summary_{TARGET_KEY}.csv"
summary_bytes = summary_path.read_bytes() summary_bytes = summary_path.read_bytes()
self.assertFalse(summary_bytes.startswith(b"\xef\xbb\xbf")) self.assertTrue(summary_bytes.startswith(b"\xef\xbb\xbf"))
self.assertIn("小区".encode("gbk"), summary_bytes) with summary_path.open(encoding="utf-8-sig", newline="") as file:
with summary_path.open(encoding="gbk", newline="") as file:
rows = list(csv.DictReader(file)) rows = list(csv.DictReader(file))
self.assertEqual(len(rows), 7) self.assertEqual(len(rows), 7)
self.assertEqual( self.assertEqual(
@@ -246,8 +245,12 @@ class PipelineTest(unittest.TestCase):
self.assertEqual(rows[1]["prev_interference_dbm"], "") self.assertEqual(rows[1]["prev_interference_dbm"], "")
self.assertEqual(rows[1]["prev_nearby_count"], "") self.assertEqual(rows[1]["prev_nearby_count"], "")
def test_gbk_csv_normalizes_non_breaking_spaces(self) -> None: def test_gbk_history_csv_normalizes_non_breaking_spaces(self) -> None:
payload = main.dict_csv_bytes(("cell_name",), [{"cell_name": "测试\u00a0小区"}]) payload = main.dict_csv_bytes(
("cell_name",),
[{"cell_name": "测试\u00a0小区"}],
main.HISTORY_CSV_ENCODING,
)
self.assertEqual(payload.decode("gbk"), "cell_name\r\n测试 小区\r\n") self.assertEqual(payload.decode("gbk"), "cell_name\r\n测试 小区\r\n")