Files
InterferenceETL/docs/assumptions.md
T

29 lines
2.3 KiB
Markdown

# Assumptions and pending decisions
## Implemented assumptions
- The source directory contains seven expected source types for each complete hour.
- The newest usable hour is the newest 16-digit filename window for which all seven types exist. A newer incomplete hour is logged and skipped.
- Each ZIP contains exactly one XLSX member. The member is selected by extension because the two 700M ZIP series use a legacy filename encoding.
- Only `Sheet0` is converted and merged. The `指标(计数器)` sheet is metadata and is not included in the summary.
- LTE/SDR CGI is derived as `{PLMN}-{eNodeBId}-{cellId}` with default PLMN `460-00`.
- NR CGI uses the source `masterOperatorId` unchanged. Confirm whether the final database should retain this source format or normalize it.
- Output CSV files use UTF-8 with BOM so they open correctly in Excel.
- Source files are read-only. The script writes only below its configured output directory.
- MySQL is the scheduled-run output. The table uses one `metric_time DATETIME` column containing the source KPI start time.
- A successful run transactionally refreshes the selected hour and deletes rows for all other hours. A selected hour older than the newest database hour is rejected to prevent data rollback.
## Pending user decisions
- Target database host, database name, user, and credentials. The default table name is `interference_hourly_summary` and can be overridden.
- Whether the seven converted full CSV files must be retained after successful database import.
- Whether the `指标(计数器)` metadata sheet must also be stored.
- Final CGI normalization rules, especially for NR `masterOperatorId`.
- Schedule minute and expected source-file arrival delay. A safe initial suggestion is hourly at minute 20.
- Historical backfill range and retention policy for generated CSV files.
- Whether an incomplete newest hour should only warn, fail the run, or wait and retry.
## Mock boundary
Automated tests generate seven in-memory XLSX/ZIP sources plus a newer incomplete hour. The mock verifies complete-hour fallback, schema validation, CSV conversion, CGI extraction, summary merging, cross-midnight windows, transactional latest-hour replacement, and database rollback protection. No mock credentials or fake database writes are present in production code.