life-echo

Author	SHA1	Message	Date
Kevin	ac49bc7f23	feat(eval): memoir A/B chapter judging and eval-web parity with dialogue - Judge baseline excerpt and library chapter separately; build_memoir_compare_summary for gate, nine-dim and leaf deltas. - Memoir SSE chapter payload: baseline_judge, compare_summary, baseline_judge_error. - MemoirJudgeOutput: loose score coercion and post-validate clamp; memoir judge prompt caps from settings. - app-eval-web: two-column MemoirScoreCard layout, MemoirCompareSummary, chapter blocks and CSS. - Add memoir_compare_summary, log_events, celery_log_context, memoir_pipeline_progress; tests and migration 0014. - Misc: memory/evidence and enrichment paths, task/orchestrator updates, internal-eval docs, env examples.	2026-04-10 10:25:15 +08:00
Kevin	b0251e5b26	feat(eval): server-side replay/phase1 timing + memoir phase1 batch chunking - Replay and memoir-submit responses include started/finished UTC and elapsed_ms; Phase1 poll exposes Redis-backed submit time and elapsed_ms_since_submit. - Phase1 batch LLM splits segments by memoir_phase1_batch_llm_chunk_size with bisect fallback per chunk; Playground shows server timings. Made-with: Cursor	2026-04-09 13:39:04 +08:00
Kevin	064ad2161d	refactor(eval+memoir)：精简内部评测路由与服务，composite/对话摘要与 judge 能力补强 - 访谈：新增 interview_state_hints，联动 orchestrator 与提示词 - 回忆录：story_pipeline_sync/state/memory/post_commit 与 Celery 任务调整 - 基建：开发用 celery broker、compose/development 脚本、依赖注入 - eval-web：移除数据集/实验/版本等页面与流式轮询，突出 Playground - 文档与单测同步	2026-04-08 21:36:12 +08:00
Kevin	78b61c076e	feat(eval): Playground GLM 评分落库并可恢复在 conversations 表增加 playground_conversation_judge_json，流式/非流式对话评审结束后写入最近一次快照（整体分、逐轮分、对比文案、错误与基线文件名等）。新增只读 GET 供前端按会话拉取；评测台 Playground 切换会话时自动恢复，并提示基线是否和当时一致。	2026-04-08 16:51:08 +08:00
Kevin	6772e1269c	feat(evaluation): memoir readiness, judge/replay updates, eval web playground Add memoir_readiness_service and router tests; extend judge schemas/services, replay_service, and conversation rubric; align story route agent, payload, prompts, and story_pipeline_sync; update agent logging, config, and DI. Document internal-eval; add replayDraft util and PlaygroundPage changes in app-eval-web.	2026-04-08 09:43:34 +08:00
Kevin	99543d04c6	feat(eval): internal-eval stack, judge fixes, and eval web overhaul - Merge internal-eval into development.sh (single Celery/infra); internal-eval.sh wraps with LIFE_ECHO_WITH_INTERNAL_EVAL; EVAL_ATTACH_ONLY for attaching 8001 when :8000 is already up; document in api/docs/internal-eval.md. - Evaluation: transcript_for_judge, judge error surfacing, rubric/schema tweaks, execution_service and router updates; tests for judge and composite eval. - Memory: ingest nested transaction for embedding/enrichment rollback safety. - Conversation WS: logger.exception for pipeline errors (avoid loguru KeyError). - app-eval-web: Playground saved replays, dialogue turns helper, hash user_id for Memoir; Memoir chapter baseline↔DB row compare with title heuristics; Stories page (#memoir-stories); Markdown + copy buttons; toolbar/panel UI; react-markdown; development proxy and fixture updates.	2026-04-07 17:18:47 +08:00
Kevin	29dec8fe32	feat/ eval	2026-04-06 23:19:20 +08:00
Kevin	ca8bcc8489	feat(evaluation): session catalog, user export import, and eval web UI - Extend evaluation API: schemas, router, repo, admin and execution services - Improve user export markdown importer; add fixtures and importer tests - Session catalog repo/service updates; internal app wiring and docs - Add internal-eval.sh helper; refresh app-eval-web (App, styles, Vite)	2026-04-06 13:49:28 +08:00
Kevin	b75edacb5f	feat/ 导出开发容器内的数据用于评估	2026-04-03 14:44:46 +08:00

9 Commits