diff --git a/README.md b/README.md index 42deb69..2b0b72c 100644 --- a/README.md +++ b/README.md @@ -25,7 +25,7 @@ For every processed group, the script also reads the latest filename-dated XLSX The summary columns are: ```text -metric_time,network_type,cgi,cell_name,interference_dbm,prev_interference_dbm,longitude,latitude,azimuth,nearby_count,prev_nearby_count +metric_time,network_type,cgi,cell_name,interference_dbm,prev_interference_dbm,longitude,latitude,azimuth,nearby_count,nearby_26g,nearby_700m,nearby_tdd,nearby_fdd,prev_nearby_count,prev_nearby_26g,prev_nearby_700m,prev_nearby_tdd,prev_nearby_fdd ``` All selected workbook rows must belong to the selected natural 15-minute group. Every summary row receives the same normalized `metric_time`; a row outside the group fails the run before database or history output. @@ -34,15 +34,17 @@ All selected workbook rows must belong to the selected natural 15-minute group. `nearby_count` is the number of other high-interference cells from the same selected time group within 1 km, across all network types. The cell itself is excluded, the 1 km boundary is included, and rows without coordinates use `0`. Candidate cells are found through a 1 km spatial grid index and confirmed with exact Haversine distance. +`nearby_26g`, `nearby_700m`, `nearby_tdd`, and `nearby_fdd` split the same neighbors by the neighbor's `network_type`. Each field defaults to `0`, and the four values always add up to `nearby_count`. + When multiple retained source rows have the same CGI, the summary keeps the row with the numerically largest interference value. Equal values keep the first row encountered. Deduplication happens before nearby-cell counting and database insertion; converted source CSVs remain complete. -`prev_interference_dbm` and `prev_nearby_count` come from the previously retained database time, matched by CGI before the current batch replaces old rows. They represent the same cell's previous-period interference value and nearby high-interference count. On the first run, or when a CGI did not exist in the previous period, both fields are empty in CSV/history output and `NULL` in MySQL. +`prev_interference_dbm`, `prev_nearby_count`, and the four `prev_nearby_26g/700m/tdd/fdd` fields come from the previously retained database time, matched by CGI before the current batch replaces old rows. They represent the same cell's previous-period interference value and nearby high-interference counts. On the first run, or when a CGI did not exist in the previous period, these fields are empty in CSV/history output and `NULL` in MySQL. The script never modifies or deletes source storage files. Before downloading source ZIP files or CellData, it compares the normalized target time with the database. A target already present, or older than the database, is not processed again. -The database table contains normalized `metric_time DATETIME`, `network_type VARCHAR(16)`, `interference_dbm DECIMAL(10,3)`, nullable `prev_interference_dbm DECIMAL(10,3)`, nullable `longitude` and `latitude`, `azimuth DECIMAL(6,2) NOT NULL DEFAULT 0`, `nearby_count INT NOT NULL DEFAULT 0`, and nullable `prev_nearby_count INT`. A successful transaction replaces the target batch and deletes every other database time, so the table retains only the latest processed group. The previous-period values are read before this replacement transaction. +The database table contains normalized `metric_time DATETIME`, `network_type VARCHAR(16)`, `interference_dbm DECIMAL(10,3)`, nullable `prev_interference_dbm DECIMAL(10,3)`, nullable `longitude` and `latitude`, `azimuth DECIMAL(6,2) NOT NULL DEFAULT 0`, the five `NOT NULL DEFAULT 0` INT count columns (`nearby_count`, `nearby_26g`, `nearby_700m`, `nearby_tdd`, `nearby_fdd`), and the five nullable INT previous-period count columns (`prev_nearby_count`, `prev_nearby_26g`, `prev_nearby_700m`, `prev_nearby_tdd`, `prev_nearby_fdd`). A successful transaction replaces the target batch and deletes every other database time, so the table retains only the latest processed group. The previous-period values are read before this replacement transaction. -After the database succeeds, the same eleven result columns are uploaded as GBK CSV to: +After the database succeeds, the same nineteen result columns are uploaded as GBK CSV to: ```text /网优日常优化数据文档/(勿删)干扰定时小时指标/干扰历史数据/YYYY-MM-DD/干扰数据处理结果_YYYYMMDDHHMMSS.csv diff --git a/aidocs/project_context.md b/aidocs/project_context.md index 66943ae..c902ae9 100644 --- a/aidocs/project_context.md +++ b/aidocs/project_context.md @@ -109,6 +109,12 @@ - The current `2026-08-07 14:00:00` database result was exported directly through the history-only path and replaced in Storage as a verified GBK CSV: 1,082 rows, 150,251 bytes, no UTF-8 BOM, and the expected eleven columns. The temporary UTF-8 backup was deleted after validation. - The newer `2026-08-07 15:00:00` source group currently contains duplicate CGI rows, for example `460-00-122737-22`. Its database transaction rolled back on the existing `(metric_time, cgi)` primary key, so the retained `14:00` database and history result were not replaced. Duplicate-row selection requires a separate business rule. +## 2026-08-11: Per-network-type nearby counts + +- Summary CSV, database table, and history CSV add `nearby_26g`, `nearby_700m`, `nearby_tdd`, and `nearby_fdd`: the same 1 km high-interference neighbors split by the neighbor's `network_type`. Each defaults to `0` and the four values always sum to `nearby_count`. +- Previous-period fields `prev_nearby_26g`, `prev_nearby_700m`, `prev_nearby_tdd`, and `prev_nearby_fdd` follow the `prev_nearby_count` rules: matched by CGI from the previously retained database time, empty/`NULL` when unmatched or on the first run. On the first run after this upgrade the previous rows only have `NULL` in the new columns, so the previous split fields stay empty for that one period and self-heal afterwards. +- Existing tables migrate automatically: current count columns are `INT NOT NULL DEFAULT 0`, previous count columns are `INT NULL`. The history export SELECT is now generated from `SUMMARY_HEADER`, so result columns stay in one place. Twenty-two tests pass. + ## 2026-08-11: Network type relabeled to 2.6G/700M/TDD/FDD - `network_type` keeps its column name but now labels radio type per source prefix: `5G干扰监控` is `2.6G` (the user treats 2.6G as the unambiguous 5G label because 700M is also 5G), `700M干扰监控` is `700M`, `SDR_TDD干扰监控` and `反开RD干扰监控` are `TDD`, and the three FDD-schema sources (`SDR_FDD干扰监控`, `5G下FDD干扰监控`, `700M下FDD干扰监控`) are `FDD`. The old `2.6G/700M/4G` grouping is gone. diff --git a/main.py b/main.py index 38db6f2..f8d11bc 100644 --- a/main.py +++ b/main.py @@ -72,6 +72,13 @@ MIN_INTERFERENCE_BY_SOURCE = { EARTH_RADIUS_KM = 6371.0088 NEARBY_RADIUS_KM = 1.0 +NEARBY_COLUMN_BY_NETWORK_TYPE = { + "2.6G": "nearby_26g", + "700M": "nearby_700m", + "TDD": "nearby_tdd", + "FDD": "nearby_fdd", +} + DEFAULT_STORAGE_ID = "stg_4d9a910d72" DEFAULT_ROOT = "/网优日常优化数据文档/(勿删)干扰定时小时指标" CELL_DATA_DIRECTORIES = ( @@ -192,7 +199,15 @@ SUMMARY_HEADER = ( "latitude", "azimuth", "nearby_count", + "nearby_26g", + "nearby_700m", + "nearby_tdd", + "nearby_fdd", "prev_nearby_count", + "prev_nearby_26g", + "prev_nearby_700m", + "prev_nearby_tdd", + "prev_nearby_fdd", ) class ProcessingError(RuntimeError): @@ -358,7 +373,7 @@ class ApiSummaryStore: START TRANSACTION; DELETE FROM `{self.table}` WHERE metric_time = {metric_literal}; INSERT INTO `{self.table}` - (metric_time, network_type, cgi, cell_name, interference_dbm, prev_interference_dbm, longitude, latitude, azimuth, nearby_count, prev_nearby_count) + (metric_time, network_type, cgi, cell_name, interference_dbm, prev_interference_dbm, longitude, latitude, azimuth, nearby_count, nearby_26g, nearby_700m, nearby_tdd, nearby_fdd, prev_nearby_count, prev_nearby_26g, prev_nearby_700m, prev_nearby_tdd, prev_nearby_fdd) VALUES {values}; DELETE FROM `{self.table}` WHERE metric_time <> {metric_literal}; @@ -392,8 +407,7 @@ class ApiSummaryStore: self._ensure_schema() metric_literal = f"'{metric_time:%Y-%m-%d %H:%M:%S}'" sql = ( - "SELECT metric_time, network_type, cgi, cell_name, interference_dbm, " - "prev_interference_dbm, longitude, latitude, azimuth, nearby_count, prev_nearby_count " + f"SELECT {', '.join(SUMMARY_HEADER)} " f"FROM `{self.table}` " f"WHERE metric_time = {metric_literal} ORDER BY cgi, cell_name" ) @@ -431,7 +445,15 @@ class ApiSummaryStore: latitude DECIMAL(10,6) NULL, azimuth DECIMAL(6,2) NOT NULL DEFAULT 0, nearby_count INT NOT NULL DEFAULT 0, + nearby_26g INT NOT NULL DEFAULT 0, + nearby_700m INT NOT NULL DEFAULT 0, + nearby_tdd INT NOT NULL DEFAULT 0, + nearby_fdd INT NOT NULL DEFAULT 0, prev_nearby_count INT NULL, + prev_nearby_26g INT NULL, + prev_nearby_700m INT NULL, + prev_nearby_tdd INT NULL, + prev_nearby_fdd INT NULL, PRIMARY KEY (metric_time, cgi) ) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 """, @@ -451,7 +473,15 @@ class ApiSummaryStore: "latitude": "DECIMAL(10,6) NULL", "azimuth": "DECIMAL(6,2) NOT NULL DEFAULT 0", "nearby_count": "INT NOT NULL DEFAULT 0", + "nearby_26g": "INT NOT NULL DEFAULT 0", + "nearby_700m": "INT NOT NULL DEFAULT 0", + "nearby_tdd": "INT NOT NULL DEFAULT 0", + "nearby_fdd": "INT NOT NULL DEFAULT 0", "prev_nearby_count": "INT NULL", + "prev_nearby_26g": "INT NULL", + "prev_nearby_700m": "INT NULL", + "prev_nearby_tdd": "INT NULL", + "prev_nearby_fdd": "INT NULL", } for column, definition in migrations.items(): if column not in existing_columns: @@ -487,7 +517,15 @@ class ApiSummaryStore: sql_decimal_literal(row["latitude"], "latitude", nullable=True), sql_decimal_literal(row["azimuth"], "azimuth"), str(int(row["nearby_count"])), + str(int(row["nearby_26g"])), + str(int(row["nearby_700m"])), + str(int(row["nearby_tdd"])), + str(int(row["nearby_fdd"])), sql_decimal_literal(row["prev_nearby_count"], "prev_nearby_count", nullable=True), + sql_decimal_literal(row["prev_nearby_26g"], "prev_nearby_26g", nullable=True), + sql_decimal_literal(row["prev_nearby_700m"], "prev_nearby_700m", nullable=True), + sql_decimal_literal(row["prev_nearby_tdd"], "prev_nearby_tdd", nullable=True), + sql_decimal_literal(row["prev_nearby_fdd"], "prev_nearby_fdd", nullable=True), ) ) + ")" @@ -966,7 +1004,7 @@ def summary_record( interference = required(record, "载波平均噪声干扰(dBm)", source_path, row_number) longitude, latitude, azimuth = (cell_metadata or {}).get(cgi, ("", "", "0")) - return { + record_out = { "metric_time": metric_time.strftime("%Y-%m-%d %H:%M:%S"), "network_type": NETWORK_TYPE_BY_SOURCE[source_type], "cgi": cgi, @@ -979,6 +1017,10 @@ def summary_record( "nearby_count": "0", "prev_nearby_count": "", } + for column in NEARBY_COLUMN_BY_NETWORK_TYPE.values(): + record_out[column] = "0" + record_out[f"prev_{column}"] = "" + return record_out def passes_interference_threshold(record: dict[str, str], source_type: str) -> bool: @@ -992,10 +1034,9 @@ def passes_interference_threshold(record: dict[str, str], source_type: str) -> b def populate_nearby_counts(rows: list[dict[str, str]]) -> None: - counts = [0] * len(rows) + type_counts = [dict.fromkeys(NEARBY_COLUMN_BY_NETWORK_TYPE.values(), 0) for _ in rows] spatial_index: dict[tuple[int, int, int], list[tuple[int, float, float]]] = {} for index, row in enumerate(rows): - row["nearby_count"] = "0" if not row["longitude"] or not row["latitude"]: continue try: @@ -1006,6 +1047,7 @@ def populate_nearby_counts(rows: list[dict[str, str]]) -> None: if not math.isfinite(longitude) or not math.isfinite(latitude) or not -180 <= longitude <= 180 or not -90 <= latitude <= 90: raise ProcessingError(f"Invalid coordinates for CGI {row['cgi']}") + column = NEARBY_COLUMN_BY_NETWORK_TYPE[row["network_type"]] bucket = spatial_bucket(longitude, latitude) for x_offset in (-1, 0, 1): for y_offset in (-1, 0, 1): @@ -1013,12 +1055,14 @@ def populate_nearby_counts(rows: list[dict[str, str]]) -> None: nearby_bucket = (bucket[0] + x_offset, bucket[1] + y_offset, bucket[2] + z_offset) for other_index, other_longitude, other_latitude in spatial_index.get(nearby_bucket, ()): if haversine_distance_km(longitude, latitude, other_longitude, other_latitude) <= NEARBY_RADIUS_KM: - counts[index] += 1 - counts[other_index] += 1 + type_counts[index][NEARBY_COLUMN_BY_NETWORK_TYPE[rows[other_index]["network_type"]]] += 1 + type_counts[other_index][column] += 1 spatial_index.setdefault(bucket, []).append((index, longitude, latitude)) - for index, count in enumerate(counts): - rows[index]["nearby_count"] = str(count) + for row, counts in zip(rows, type_counts, strict=True): + for column, count in counts.items(): + row[column] = str(count) + row["nearby_count"] = str(sum(counts.values())) def deduplicate_by_cgi(rows: list[dict[str, str]]) -> int: @@ -1038,14 +1082,11 @@ def deduplicate_by_cgi(rows: list[dict[str, str]]) -> int: def populate_previous_values(rows: list[dict[str, str]], previous_rows: list[dict[str, str]]) -> None: previous_by_cgi = {row["cgi"]: row for row in previous_rows} + copied_columns = ("interference_dbm", "nearby_count", *NEARBY_COLUMN_BY_NETWORK_TYPE.values()) for row in rows: previous = previous_by_cgi.get(row["cgi"]) - if previous is None: - row["prev_interference_dbm"] = "" - row["prev_nearby_count"] = "" - continue - row["prev_interference_dbm"] = previous["interference_dbm"] - row["prev_nearby_count"] = previous["nearby_count"] + for column in copied_columns: + row[f"prev_{column}"] = previous[column] if previous is not None else "" def spatial_bucket(longitude: float, latitude: float) -> tuple[int, int, int]: diff --git a/tests/test_pipeline.py b/tests/test_pipeline.py index 80e343f..50e8935 100644 --- a/tests/test_pipeline.py +++ b/tests/test_pipeline.py @@ -57,7 +57,15 @@ class PipelineTest(unittest.TestCase): "latitude", "azimuth", "nearby_count", + "nearby_26g", + "nearby_700m", + "nearby_tdd", + "nearby_fdd", "prev_nearby_count", + "prev_nearby_26g", + "prev_nearby_700m", + "prev_nearby_tdd", + "prev_nearby_fdd", ], ) rows_by_cgi = {row["cgi"]: row for row in rows} @@ -220,16 +228,34 @@ class PipelineTest(unittest.TestCase): def test_nearby_count_uses_high_interference_rows_with_coordinates(self) -> None: rows = [ - nearby_row("a", "113.000000", "22.000000"), - nearby_row("b", "113.005000", "22.000000"), - nearby_row("c", "113.020000", "22.000000"), - nearby_row("d", "113.000000", "22.000000"), - nearby_row("e", "", ""), + nearby_row("a", "113.000000", "22.000000", "2.6G"), + nearby_row("b", "113.005000", "22.000000", "700M"), + nearby_row("c", "113.020000", "22.000000", "TDD"), + nearby_row("d", "113.000000", "22.000000", "FDD"), + nearby_row("e", "", "", "TDD"), ] main.populate_nearby_counts(rows) self.assertEqual([row["nearby_count"] for row in rows], ["2", "2", "0", "2", "0"]) + by_cgi = {row["cgi"]: row for row in rows} + self.assertEqual( + [by_cgi["a"][column] for column in ("nearby_26g", "nearby_700m", "nearby_tdd", "nearby_fdd")], + ["0", "1", "0", "1"], + ) + self.assertEqual( + [by_cgi["b"][column] for column in ("nearby_26g", "nearby_700m", "nearby_tdd", "nearby_fdd")], + ["1", "0", "0", "1"], + ) + self.assertEqual( + [by_cgi["d"][column] for column in ("nearby_26g", "nearby_700m", "nearby_tdd", "nearby_fdd")], + ["1", "1", "0", "0"], + ) + for row in rows: + per_type_total = sum( + int(row[column]) for column in ("nearby_26g", "nearby_700m", "nearby_tdd", "nearby_fdd") + ) + self.assertEqual(per_type_total, int(row["nearby_count"])) self.assertLess(main.haversine_distance_km(113.0, 22.0, 113.005, 22.0), 1.0) self.assertGreater(main.haversine_distance_km(113.0, 22.0, 113.02, 22.0), 1.0) @@ -241,13 +267,25 @@ class PipelineTest(unittest.TestCase): previous["cgi"] = "matched" previous["interference_dbm"] = "-106.25" previous["nearby_count"] = "8" + previous["nearby_26g"] = "3" + previous["nearby_700m"] = "2" + previous["nearby_tdd"] = "2" + previous["nearby_fdd"] = "1" main.populate_previous_values(rows, [previous]) self.assertEqual(rows[0]["prev_interference_dbm"], "-106.25") self.assertEqual(rows[0]["prev_nearby_count"], "8") + self.assertEqual(rows[0]["prev_nearby_26g"], "3") + self.assertEqual(rows[0]["prev_nearby_700m"], "2") + self.assertEqual(rows[0]["prev_nearby_tdd"], "2") + self.assertEqual(rows[0]["prev_nearby_fdd"], "1") self.assertEqual(rows[1]["prev_interference_dbm"], "") self.assertEqual(rows[1]["prev_nearby_count"], "") + self.assertEqual(rows[1]["prev_nearby_26g"], "") + self.assertEqual(rows[1]["prev_nearby_700m"], "") + self.assertEqual(rows[1]["prev_nearby_tdd"], "") + self.assertEqual(rows[1]["prev_nearby_fdd"], "") def test_duplicate_cgi_keeps_maximum_interference(self) -> None: rows = [database_row(), database_row(), database_row()] @@ -292,12 +330,15 @@ class PipelineTest(unittest.TestCase): self.assertTrue(any("ADD COLUMN `latitude`" in query for query in queries)) self.assertTrue(any("ADD COLUMN `azimuth`" in query for query in queries)) self.assertTrue(any("ADD COLUMN `nearby_count`" in query for query in queries)) + for column in ("nearby_26g", "nearby_700m", "nearby_tdd", "nearby_fdd"): + self.assertTrue(any(f"ADD COLUMN `{column}`" in query for query in queries)) + self.assertTrue(any(f"ADD COLUMN `prev_{column}`" in query for query in queries)) self.assertTrue(any("ADD COLUMN `prev_interference_dbm`" in query for query in queries)) self.assertTrue(any("ADD COLUMN `prev_nearby_count`" in query for query in queries)) script = next(payload for endpoint, payload in client.posts if endpoint.endswith("/run-script")) self.assertTrue(script["single_session"]) self.assertIn("START TRANSACTION", script["content"]) - self.assertIn("NULL, NULL, NULL, 0, 0, NULL", script["content"]) + self.assertIn("NULL, NULL, NULL, 0, 0, 0, 0, 0, 0, NULL, NULL, NULL, NULL, NULL", script["content"]) self.assertIn("CONVERT(0x", script["content"]) self.assertNotIn("source_type", script["content"]) self.assertNotIn("source_path", script["content"]) @@ -507,15 +548,24 @@ def database_row() -> dict[str, str]: "latitude": "", "azimuth": "0", "nearby_count": "0", + "nearby_26g": "0", + "nearby_700m": "0", + "nearby_tdd": "0", + "nearby_fdd": "0", "prev_nearby_count": "", + "prev_nearby_26g": "", + "prev_nearby_700m": "", + "prev_nearby_tdd": "", + "prev_nearby_fdd": "", } -def nearby_row(cgi: str, longitude: str, latitude: str) -> dict[str, str]: +def nearby_row(cgi: str, longitude: str, latitude: str, network_type: str = "2.6G") -> dict[str, str]: row = database_row() row["cgi"] = cgi row["longitude"] = longitude row["latitude"] = latitude + row["network_type"] = network_type return row