Close the Learning Loop
Keep linked work history and memory; curate supported findings for later work.
An accepted outcome leaves work history and supported findings. Preserve them with their source owners, then keep those records portable across the runtimes considered in Factor X.
The rule
Store enough context to resume work and retrieve relevant evidence later.
Work history, project facts and reviewed knowledge serve different needs; link them without building a second authority for status.
Beads provides persistent work items, dependencies, notes and history, plus a small project-hint memory store through bd remember, bd memories and bd recall. It is more than a queue. Linked intent, decisions and results make earlier work retrievable across sessions. Larger reports and source artifacts can stay in their own stores with links from the work item.
AgentOps Memory can find relevant context, draft a report from supported evidence or help curate project knowledge. Private drafts need appropriate storage and support and disclosure review before import into shared documents. Recall is evidence to evaluate, not an instruction to obey.
The evidencesupported
Lab studies show memory helps agents when its quality is controlled and hurts when it is not. Whether a team's store pays off over months is close to untested.
- Xiong et al., 2025 found agents closely repeat what a similar stored record did, so an error kept in memory carries into later tasks. Admitting only results checked against ground truth scored best on all four agents tested. An off-the-shelf LLM judge as the gate gave mixed results, sometimes below a memory that never changed.
- ReasoningBank, 2025 stored strategies distilled from both successes and failures. It outperformed memories of raw trajectories on web browsing and software engineering benchmarks.
- ACE, 2025 found that repeatedly rewriting an agent's stored context erodes its detail. Structured, incremental updates kept the detail and improved agent benchmark results by 10.6%.
Helwig, 2026 offers what its author believes is the first months-long instrumented record of an agent memory system in development use. It states its limits plainly: one project, no control arm, self-reported.
Put it to work
- Record decisions, outcomes, unresolved questions and evidence links in the work item.
- Keep small project hints in native memory with enough context to interpret them.
- Curate reusable findings only when a consumer needs them.
- Recheck applicability and freshness when recalled context affects a decision.
- Add to shared memory in small, sourced entries; do not let an agent rewrite the whole store.
- Correct or drop an entry when later work shows it misled.
Example · the monthly operations report
The report's Beads item links its brief, ledger snapshot, checks and approved output. A recalled reversal rule guides next month's investigation. Memory can assemble a cited handoff or report from those records without treating last month's numbers as current.
Failure signal
- Useful decisions exist only in a lost transcript.
- A retrieved lesson has no supporting source or applicability boundary.
- The system claims improvement without evidence from later work.
Done looks like
Later work can recover the relevant intent, evidence and decision. Any reusable guidance has support, an owner and a clear scope.