feat: v0.3 - Manus replay + token tracking + SIGTERM fix + accelerate button
This is the v0.3 milestone commit before the v0.4 big version push.
Major themes: process replay, runtime stability, cost observability.
## New Features
- **Manus-style process replay** (frontend + backend)
- `GET /api/simulation/<id>/replay` returns full workflow + agents + rounds + aggregate
- `frontend/src/views/SimulationReplayView.vue` 3-column layout (workflow / actions / stats)
- bottom scrubber with play/pause/step + 5 speed levels (0.5x-10x)
- filters out stale actions from previous runs via latest simulation_start timestamp
- **Token usage tracking** (`backend/app/utils/token_tracker.py`)
- process-wide stage→model→tokens counter
- LLMClient auto-records prompt/completion tokens after each call
- stages tagged at API entry: step1_ontology, step2_graph_build, step3_prepare, step5_report
- `GET /api/usage/summary` for live stats + CNY cost estimate
- `GET /api/usage/estimate-simulation` for OASIS subprocess estimation
- pricing table for GLM/SiliconFlow/MiniMax/OpenAI/Anthropic models
- documented as internal-use, removed from customer-facing builds
- **Step 2 "Skip & Continue" button**
- lets user stop profile generation early and proceed with what's already generated
- `simulation_manager.request_accelerate()` + cancel_check in oasis_profile_generator
- new endpoint `POST /api/simulation/prepare/accelerate`
## Critical Bug Fix
- **SIGTERM no longer kills running simulation subprocess**
- root cause: `SimulationRunner.register_cleanup()` registered SIGTERM/SIGINT/SIGHUP handlers
that called `os.killpg` on every tracked sim child, even though spawn already used
`start_new_session=True` to give children isolated sessions
- fix: neutered `register_cleanup` to a no-op; `cleanup_all_simulations` itself preserved
for explicit stop_simulation paths
- validated: killed Flask backend twice, simulation subprocess kept running
- impact: hot-reload backend code without interrupting in-flight simulations
## Performance & Tuning
- semaphore 30 → 100 (twitter + reddit) for higher LLM concurrency
- discovered 200 agents as memory/cost/statistical sweet spot for 8G server
(503 agents OOMs both platforms; 200 agents fits cleanly with 95% confidence margin)
## Documentation
- **PRD.md** rewritten as v0.3 baseline (10 chapters + 2 appendices, 639 lines)
- product positioning across 3 usage modes (one-shot / model-reuse / SaaS)
- v0.4 roadmap: domestic platforms (douyin/wechat/xiaohongshu/weibo), fork sim, multi-tenant
- operational lessons: HF mirror, Tencent PyPI mirror, GLM-4-Flash choice, agent count
- decision log with dates
- per-stage token/cost breakdown for typical 200-agent run
- **README.md / README-ZH.md** updated with replay step + Graphiti+Neo4j
## Files Touched
19 files changed, 1105 insertions(+), 246 deletions(-)