ForecastOps is a local-first observability and evaluation layer for production forecasts. It works with the forecasting code you already have.
Add one line after .predict(), then run fops ui.
import forecastops as fops
forecast = model.predict(future)
run = fops.capture(
forecast,
project="site-traffic",
series_id="homepage",
cutoff=train_df["ds"].max(),
actuals=actuals_df,
)fops uiForecastOps stores forecast artifacts locally as Parquet, writes run metadata to DuckDB, computes horizon-aware metrics, generates static HTML reports, and serves a read-only local UI. It does not train forecasting models, require a cloud account, or upload raw forecast data.
From PyPI:
pip install forecastopsFrom source:
git clone https://github.com/Parisi-Labs/forecastops.git
cd forecastops
pip install -e .python -m venv .venv
source .venv/bin/activate
pip install -e .
python examples/generic_dataframe.py
fops report --latest
fops uiOpen http://127.0.0.1:4784 after starting the UI.
fops ui serves a read-only explorer for the local store:
- Runs — every captured run with horizon, points, MAE, WAPE, bias, coverage, coverage gap, skill, and validation status; filterable and sortable.
- Run detail — headline metrics, a forecast inspector with one chart per series, and a diagnostics cockpit: residual distribution, error by horizon, per-series worst offenders, and per-regime breakdowns — plus metrics, validation, residuals, artifacts, and the capture trace timeline.
- Projects — runs grouped by project with error trends across captures.
- Groups — experiment and backtest groups with run counts and mean error; open a group to see per-metric mean ± std and stability across its runs.
- Compare — metric deltas and regressions between any two runs, backed
by
fops diff.
capture: normalize forecasts from existing workflows.ForecastSchema: map arbitrary dataframe columns to canonical semantics.validate: catch schema, timestamp, duplicate, interval, and leakage issues.evaluate: compute MAE, RMSE, WAPE, sMAPE, bias, coverage, coverage gap, interval width, pinball loss (for quantile forecasts), and count — sliced by horizon and by any categorical columns you keep (e.g. region, holiday_flag, event_type).compare: calculate benchmark metrics and skill.backtest: evaluate a rolling-origin forecast panel as one grouped run set, with per-cutoff and aggregate (mean/std) metrics.diff: compare two captured runs.diagnose: a machine-readable diagnosis of a run — overall metrics, skill, worst horizons/series/regimes, validation, and artifact URIs — for agents and scripts (fops diagnose <run_id>).- groups: tag related runs with
capture(group=...)(or abacktest) and browse them together in the UI. - local store:
.forecastops/forecastops.duckdbplus Parquet artifacts. - UI: local read-only browser explorer for runs, metrics, residuals, validation, artifacts, and run differences.
ForecastOps stores metric values as machine-readable ratios or forecast-unit values, not display-formatted percentages:
- MAE and RMSE are in the same units as the forecast target.
- WAPE is a ratio, so
0.12means 12% weighted absolute percentage error. - sMAPE is the full symmetric MAPE ratio
2 * abs(yhat - actual) / (abs(actual) + abs(yhat)), with values in[0, 2]. - Bias is mean signed error,
mean(yhat - actual). Positive bias means the forecast overestimated actuals; negative bias means it underestimated them. - Coverage is the empirical interval hit rate. When
interval_levelis available as either a ratio (0.9) or percentage (90), ForecastOps also emitscoverage_gap = coverage - interval_level. A positive gap means overcoverage; a negative gap means undercoverage.
ForecastOps is local-first by default:
- binds the UI to
127.0.0.1and refuses other hosts unless you pass--allow-remote - makes no outbound network calls
- stores raw forecast points in the configured local store
- emits OpenTelemetry only when explicitly enabled
- avoids raw forecast points in telemetry
pip install -e ".[dev]"
pytest
ruff check .
mypy forecastopsApache-2.0. See LICENSE.