Files
InterferenceETL/docs/assumptions.md
T

27 lines
1.8 KiB
Markdown

# Assumptions and pending decisions
## Implemented assumptions
- The source directory contains seven expected source types for each complete hour.
- The newest usable hour is the newest 16-digit filename window for which all seven types exist. A newer incomplete hour is logged and skipped.
- Each ZIP contains exactly one XLSX member. The member is selected by extension because the two 700M ZIP series use a legacy filename encoding.
- Only `Sheet0` is converted and merged. The `指标(计数器)` sheet is metadata and is not included in the summary.
- LTE/SDR CGI is derived as `{PLMN}-{eNodeBId}-{cellId}` with default PLMN `460-00`.
- NR CGI uses the source `masterOperatorId` unchanged. Confirm whether the final database should retain this source format or normalize it.
- Output CSV files use UTF-8 with BOM so they open correctly in Excel.
- Source files are read-only. The script writes only below its configured output directory.
## Pending user decisions
- Target database connection, database name, table name, and credentials.
- Whether the seven converted full CSV files must be retained after successful database import.
- Whether the `指标(计数器)` metadata sheet must also be stored.
- Final CGI normalization rules, especially for NR `masterOperatorId`.
- Schedule minute and expected source-file arrival delay. A safe initial suggestion is hourly at minute 20.
- Historical backfill range and retention policy for generated CSV files.
- Whether an incomplete newest hour should only warn, fail the run, or wait and retry.
## Mock boundary
Automated tests generate seven in-memory XLSX/ZIP sources plus a newer incomplete hour. The mock verifies complete-hour fallback, schema validation, CSV conversion, CGI extraction, summary merging, and cross-midnight windows. No mock credentials or fake database writes are present in production code.