A Python script that collects environmental sensor data from the Lisbon City Council's open data portal (Lisboa Aberta) and builds a local history per station.
The Lisboa Aberta portal provides a real-time snapshot of environmental measurements (air quality, meteorology, noise) from monitoring stations across Lisbon. However, the API only returns the latest reading for each sensor — no historical data.
This script solves that by:
- Fetching the complete snapshot once per run (single HTTP request)
- Filtering for your configured stations
- Appending only new readings to per-station CSV files (deduplicated by
id+timestamp) - Running periodically (e.g., via Windows Task Scheduler), each station's CSV becomes your local historical record
- Multi-station support: Monitor any number of stations by editing
sensors.json— no code changes needed - Single API call per run: No matter how many stations you configure, the API is queried exactly once (fair use)
- Deduplication: Re-running with the same snapshot writes zero duplicate rows
- Standard library only: No external dependencies (Python 3.11+)
- UTF-8 CSV output: Opens cleanly in Excel
- Clear CLI:
--configand--out-diroptions
# 1. Clone and enter the project
git clone <your-repo-url>
cd lisboa-estacao055-avenida-brasil
# 2. Create and activate a virtual environment
python -m venv .venv
.\.venv\Scripts\Activate.ps1
# If activation fails: Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
# 3. Install test dependencies (optional, for development)
pip install -r requirements-dev.txt
# 4. Run the collector
python -m src.fetch_cml_env_sensorsOn first run, this creates CSV files in data/ — one per configured station (e.g., data/estacao_0055.csv).
Edit sensors.json to add, remove, or rename stations:
{
"output_dir": "data",
"sensors": [
{ "id": "055", "name": "Av Brasil" },
{ "id": "046", "name": "Entrecampos" },
{ "id": "049", "name": "Av Roma (cruz. Av. EUA)" },
{ "id": "073", "name": "Parque José Gomes Ferreira" },
{ "id": "037", "name": "Praça de Espanha" },
{ "id": "004", "name": "Alcântara" },
{ "id": "001", "name": "Calçada da Ajuda" },
{ "id": "068", "name": "Av. Cidade do Porto (Olivais/Aeroporto)" }
]
}id: The 2-3 digit station number (e.g.,55,055,ID55all work — normalized internally)name: Human-readable label (used in output and CSV)output_dir: Optional; defaults todata. Can be overridden with--out-dir
- Open
sensors.json - Add a new entry to the
sensorsarray:{ "id": "099", "name": "Nova Estação" } - Save and run the script — a new
data/estacao_0099.csvwill be created automatically
No code changes required.
python -m src.fetch_cml_env_sensors --config sensors.json --out-dir data| Option | Description | Default |
|---|---|---|
--config |
Path to JSON config file | sensors.json |
--out-dir |
Output directory for CSV files | Value from config's output_dir (or data) |
Each run prints a per-station summary:
[0055 - Av Brasil] API address: {'Av Brasil'}
[0055 - Av Brasil] 4 new reading(s) added to data/estacao_0055.csv
[0046 - Entrecampos] API address: {'Entrecampos'}
[0046 - Entrecampos] no new readings (data/estacao_0046.csv unchanged)
If a configured station has no matching data in the API snapshot, a warning is printed to stderr and other stations continue processing.
Each station's CSV (estacao_XXXX.csv) contains these columns:
| Column | Description |
|---|---|
id |
Full 10-char sensor ID from API (e.g., QA00NO0055) |
station_id |
4-digit station suffix (e.g., 0055) |
sensor_name |
Human-readable name from config |
parameter_code |
4-char parameter code from sensor ID (e.g., 00NO) |
address |
API-reported address for the station |
latitude |
Station latitude |
longitude |
Station longitude |
date_utc |
Reading timestamp in ISO 8601 UTC (e.g., 2024-01-15T10:30:00+00:00) |
value |
Numeric reading value |
unit |
Unit of measurement (e.g., µg/m³, °C) |
fetched_at |
When this script fetched the snapshot (ISO 8601 UTC) |
To collect data automatically, create a scheduled task:
- Open Task Scheduler → Create Basic Task
- Trigger: Daily (or hourly — match the sensor update cadence)
- Action: Start a program
- Program:
D:\path\to\.venv\Scripts\python.exe - Arguments:
-m src.fetch_cml_env_sensors --config D:\path\to\sensors.json - Start in:
D:\path\to\project
- Program:
Recommendation: Run hourly. The sensor network typically updates hourly; running more frequently wastes API quota without new data.
The project includes a Dockerfile and docker-compose.yml for containerized deployment.
# 1. Build the image
docker compose build
# 2. Run once (for testing)
docker compose run --rm lisboa-sensor-collector
# 3. For hourly runs, use TrueNAS Scale's built-in scheduler:
# - Go to Apps → Custom App → docker-compose
# - Set the command to run hourly via cron:
# docker compose -f /path/to/docker-compose.yml run --rm lisboa-sensor-collector| Method | How it works | Best for |
|---|---|---|
| Host scheduler (recommended) | TrueNAS cron runs docker compose run --rm lisboa-sensor-collector hourly |
TrueNAS Scale, any host with cron |
| Internal cron | Container runs its own cron daemon (CRON_SCHEDULE=0 * * * *) |
Standalone Docker hosts |
-
Build the image on TrueNAS:
# On TrueNAS shell or via SSH cd /mnt/pool/path/to/project docker compose build
-
Create a cron job in TrueNAS Scale:
- Go to System Settings → Advanced → Cron Jobs
- Add new cron job:
- Schedule:
0 * * * *(hourly at minute 0) - Command:
docker compose -f /mnt/pool/path/to/project/docker-compose.yml run --rm lisboa-sensor-collector - User: root (or your app user)
- Schedule:
-
Verify it works:
# Check logs docker compose -f /path/to/docker-compose.yml run --rm lisboa-sensor-collector # Check generated CSVs ls -la /mnt/pool/path/to/project/data/
Uncomment the lisboa-sensor-collector-cron service in docker-compose.yml and deploy:
# In docker-compose.yml, uncomment this service:
# lisboa-sensor-collector-cron:
# ...
# environment:
# - CRON_SCHEDULE=0 * * * *Then run:
docker compose up -d lisboa-sensor-collector-cron| Variable | Default | Description |
|---|---|---|
CONFIG_FILE |
sensors.json |
Config file path inside container |
OUT_DIR |
data |
Output directory for CSVs |
CRON_SCHEDULE |
(empty) | If set, runs cron with this schedule (e.g., 0 * * * * for hourly) |
TZ |
UTC |
Timezone for timestamps |
| Host Path | Container Path | Purpose |
|---|---|---|
./data |
/app/data |
Persistent CSV storage |
./sensors.json |
/app/sensors.json |
Config (read-only) |
For a more integrated TrueNAS experience, create a Custom App:
- Apps → Discover Apps → Custom App
- Repository:
https://github.com/syntaxterror9/cmlenvsensor.git(or local path) - Docker Compose: Use the project's
docker-compose.yml - Storage: Map
/mnt/pool/yourdataset/lisboa-sensor/data→/app/data - Environment: Set
TZ=Europe/Lisbonif you want local time in logs - Schedule: Use TrueNAS cron (System Settings → Cron Jobs) to run the collector hourly
The Dockerfile uses python:3.13-slim which supports amd64 and arm64. For TrueNAS Scale (typically amd64):
# Build specifically for amd64
docker build --platform linux/amd64 -t lisboa-sensor-collector .# Pull latest code
git pull
# Rebuild
docker compose build --no-cache
# Run
docker compose run --rm lisboa-sensor-collector# Run all unit tests (no network access)
.\.venv\Scripts\python.exe -m pytest tests/ -vAll 42 unit tests cover:
- Config loading (valid, missing, invalid JSON, empty sensors, missing id/name)
- Station suffix normalization (
55,055,ID55,id055→0055) - Date parsing (valid 18-digit, malformed, short numbers)
- Coordinate extraction (dict, string dict, missing, malformed)
- Station filtering (single/multiple stations, various ID formats)
- CSV persistence (create, append, deduplication)
- Multi-station processing (separate files, missing station warnings)
- CLI (single fetch for multiple stations, config overrides)
- Source: Lisboa Aberta / Câmara Municipal de Lisboa
- Dataset: "Monitorização de Parâmetros Ambientais da Cidade de Lisboa"
- API Endpoint:
http://opendata-cml.qart.pt:8080/lastmeasurements - License: Creative Commons Attribution 3.0 (CC-BY 3.0)
- Required attribution: "Câmara Municipal de Lisboa — Lisboa Aberta"
lisboa-estacao055-avenida-brasil/
├── CLAUDE.md # Project instructions for Claude Code
├── srs.md # Software Requirements Specification
├── README.md # This file
├── requirements-dev.txt # Test dependencies (pytest)
├── sensors.json # Station configuration (versioned)
├── .gitignore
├── src/
│ └── fetch_cml_env_sensors.py # Main script
├── tests/
│ ├── test_lisboa_estacoes.py # Unit tests
│ └── fixtures/
│ └── sensors_sample.json # Sample API response for tests
└── data/ # Generated CSVs (NOT versioned)
├── estacao_0055.csv
├── estacao_0046.csv
└── ...
.\.venv\Scripts\python.exe src/fetch_cml_env_sensors.pyEdit sensors.json, add an entry to sensors, run the script. Done.
- Python 3.11+, type hints, small functions
- Standard library only (no
requests,pandas, etc.) - Run tests before committing:
pytest tests/
| Issue | Solution |
|---|---|
ModuleNotFoundError |
Activate the venv: .\.venv\Scripts\Activate.ps1 |
PowerShell blocks Activate.ps1 |
Run Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass |
| "Config file not found" | Ensure sensors.json exists in the working directory |
| "No sensors defined" | Check sensors.json has a non-empty sensors array |
| CSV encoding issues in Excel | Files are UTF-8; use Data → From Text/CSV in Excel |
This project is for personal/educational use. The collected data is subject to the Lisboa Aberta CC-BY 3.0 license — retain attribution to Câmara Municipal de Lisboa if you publish or share the data.
This is not an official CML product. The API endpoint is public but undocumented; if it changes or disappears, the script will fail gracefully with an error message.