Dashboard
Monitor training runs in real time
trainer ships a real-time, read-only web dashboard for monitoring LLaMA-Factory (HuggingFace Trainer) training runs. It is a React frontend + a Python server.py backend under dashboard/, parses the standard training output directly, and can compare multiple runs or connect to wandb.
No public dashboard URL is checked in
The dashboard exposes run names, metrics, configuration, and logs. Keep it on localhost by default. For remote access, use an authenticated named tunnel with Cloudflare Access or an equivalent identity-aware proxy; never publish a quick tunnel URL in this repository.
Monitoring only
This dashboard monitors training — it never launches or controls it. To launch training, use the pipeline scripts. To drive training interactively, use the built-in Gradio LLaMA Board (llamafactory-cli webui), which is a separate tool.
What it monitors
For each run it reads, from artifacts/model/<run>/:
| File | Used for |
|---|---|
trainer_log.jsonl | live per-step loss, lr, grad norm, timing (primary source) |
trainer_state.json | eval points and extra metric keys from log_history |
*_results.json | final summaries (loss, runtime, samples/s, FLOPs) |
and the raw console logs from artifacts/logs/.
Run it
From subblock/trainer/:
cd dashboard
./start_dashboard.sh # builds the frontend (first run) then serves :8091Open http://localhost:8091. By default it reads ../artifacts/model (runs) and ../artifacts/logs (console logs) — the same locations the training scripts write to.
How a "run" is detected
Any immediate sub-directory of the save dir containing trainer_log.jsonl or trainer_state.json is treated as a run. Its state is running (jsonl updated < 3 min ago and percentage < 100), finished (all_results.json present or percentage ≈ 100), or unknown.
Panels
- Overview — step, epoch, progress, train/eval loss, lr, grad norm, step time, ETA
- Training — loss, learning-rate schedule, gradient norm, epoch progress
- Evaluation —
eval_lossover steps, best checkpoint (empty if no eval split) - Performance — samples/s, steps/s, runtime, total FLOPs + per-step timing
- Compare — overlay metrics from multiple runs on shared charts (PNG/CSV export)
- AI Analysis — LLM-generated diagnostic report (configure a profile in Settings)
- Logs — raw
train_*.logconsole viewer - Explorer — plot any available metric key
- Settings — LLM API profiles (stored in browser localStorage only)
The UI supports a Chinese/English toggle and dark/light themes.
Expose externally
For temporary private debugging, start_dashboard.sh can still start a quick
tunnel:
TUNNEL=true ./start_dashboard.shThe script prints a temporary URL once the tunnel is up. Do not copy that URL into docs, issues, logs, or source control.
Quick tunnels are unauthenticated
Anyone with the URL can read the dashboard (logs, metrics, run names). Treat the URL as a secret and stop the tunnel when done. For persistent, access-controlled exposure, use a named tunnel + Cloudflare Access instead.
wandb integration (optional)
Point the server at a wandb project to pull remote run metrics alongside the local files:
export WANDB_API_KEY=... WANDB_ENTITY=... WANDB_PROJECT=llama-factory
python server.py --port 8091 --save-dir ../artifacts/model \
--wandb-entity "$WANDB_ENTITY" --wandb-project "$WANDB_PROJECT" --static-dir distSee dashboard/README.md for the full CLI options and environment variables.
Two different sites
This dashboard is a live local server that reads training output. The documentation site you are reading is a separate static fumadocs site under docs/; see Build & deploy the docs site.