Skip to content

Latest commit

 

History

40 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

The research direction is:

how agents can accumulate useful experience across sessions, and where that experience should live: explicit memory, contextual lessons, executable tools, model weights, or runtime rules.

Read the executive summary for the program's main findings, their practical implications and the limits of the evidence to date.

The task is to read the previous research, perspectives and then put concepts from those files into contact through comparable tests.

The source guide links the literature review, and the study map connects mechanisms, evidence, and candidate questions. Its worked examples explain neural writes, reads, and resets. The ancillary-study approach describes how independent studies inform the root's theory and synthesis.

Contributors can use the repository list and clone instructions to retrieve all ancillary studies or select them by name.

Latest accepted publication — 20 September: Lesson acceptance at cbe62b9 finds no added benefit from paid probes once competent cheap review identifies visible SQL defects. Both curated policies complete 6/6 cases; ordinary raw reuse with repair also reaches 6/6 current outputs, while one returned program still fails on independent data. Cheaply corrected code handles all six cases without later model calls. The bounded phase closes on that explained solution; the value of evidence beyond competent ordinary review remains open. LA1–LA3 retain their original wording and now have bounded assessments. No follow-up is commissioned.

Preceding accepted publication — 18 September: Continuing consolidation at 80c17db establishes shorter familiar SQL execution after purposeful diagnosis, but its periodic adapter strategy completes 6/20 fresh jobs versus external reuse's 7/20. All arms retain the same source experience. The cumulative teaching contains only submit/finish actions, and fresh adapter runs never search history or execute saved queries. One accidental output match limits interpretation of the apparent update regression. The bounded phase is complete; useful continuing learning remains open. The review records the next research preference without starting another experiment.

Applied prototype reviewed — 19 September: Construct Runtime now includes a separate trainable memory model alongside its fixed primary, machine-facing task runtime and portable state. Two learning points improve original final completion from frozen memory's 9/28 to learned memory's 26/28; ordinary retrieval completes 28/28 at lower cost. Authority remains supplied, and changing record IDs/order breaks the learned reader. The reviewed local publication is 09f6683; root regraded 176 saved tasks and ran software checks. The derivative owns its software and experiments outside this root.

First weight-consolidation publication — 17 September: Weight consolidation completes its bounded phase at 491c0f8. With full source access, an adapter improves manual use of retained code from 3/12 to 5/12 complete tasks. A developed contract compiler brings both conditions and direct execution to 12/12; the adapter adds twelve repair calls. A corrected evidence-access flaw and all failed runs remain recorded. This is an explained limit of a finite synthetic language, not a general result against weights or a test of repeated consolidation. WC1–WC3 and their assessments remain distinct. No adapter was promoted; the completed follow-up is recorded above.

Preceding accepted publications — 17 September: Procedural migration and correction lineage complete their bounded phases. Unchanged procedural guidance and paid recipient selection both complete 12/12 fresh tasks; selection uses 83,495 tokens including setup versus 28,153 for inheritance. Guidance is useful to both tested models, but extra selection does not repay its cost at this horizon. Acquired lineage reduces recomputed fields from 48 to 18, yet whole-state rebuilding is cheaper. All four correction methods repair 6/6 states and complete 4/6 later task episodes; the shared arithmetic failures occur after correct repair. The PM1–PM3 and CL1–CL3 assessments preserve the original expectations and distinguish measured results from unresolved generalizations. No continuation is commissioned.

The preceding executable-retention assessment finds all four policies complete 8/8 fresh requests. Ready and archived source need zero future model calls, lessons-only reconstruction eight, and retaining the first reconstruction one. Source availability and persistence policy explain avoided work; acquired-lesson value and changed requirements remain untested.

The evidence-use assessment accepts its completed bounded phase. A varied-history adapter improves complete outcomes through two corrections: 65/96 uses versus the stronger retained lesson's 61/96, with 108.58 seconds of measured training plus use versus 121.16 seconds of lesson use. Further training regresses to 45/96. This is a narrow positive allocation result: cached lessons, checkpoint selection and full construction/repair costs remain untested; the supplied executable completes 96/96. Original EU1–EU3 predictions and assessments remain identifiable. No follow-up experiment is commissioned.

Current work prioritizes neural memory, live weight updates, and learned memory policies. Fifteen ancillary projects have published accepted bounded contributions. Update-source selection established reproducible local adapter updates without a task-success advantage in its final full-context comparison. Procedure acquisition and reuse found that its tested adapter transferred less reliably than retained examples and did not demonstrate acquisition-cost repayment at comparable useful accuracy. Neural memory depth (local) found that interactions among initial representations and training streams strongly affect whether a tiny deeper memory learns full or partial recall. Correcting a verified derivative omission and increasing gradient-refresh frequency did not provide a uniform remedy; the differing published depth trends in Titans and Modular TTT remain unexplained.

Procedure transfer (local) first found poor distillation acquisition and partial transfer through imitation. Its subsequent diagnosis (local) establishes a working forward-KL acquisition checkpoint and a controlled repair of failed routing from identical weights. Both selected learners route all 48 new development calls correctly, while identifier production remains unreliable. These diagnostic results explain part of the original failure without establishing a forward-KL transfer advantage. The root assessment records the evidence, checks and limits.

The third-phase root assessment, 15 September, accepts all three completed phases and evaluates the preserved M1–M3 expectations. The broader milestone is met within small controlled tasks: learned policies complete state-changing procedures through recurring learning and revision, while learned memory supports cross-event questions absent from writer training. The latest state/support assessment accepts S2's fourth phase and explains part of the maintenance divergence: the value of revision support depends on the inherited learned state. The preceding follow-up assessment records how these questions arose.

Latest publication What we learned
Lesson acceptance Paid probes add cost without benefit over competent cheap review. Raw repair fixes every current output but leaves one latent program defect; cheaply corrected code supports direct reuse.
Weight consolidation Following the finite-compiler result, a database continuation reduces familiar execution calls but adds no fresh completion or demonstrated cost repayment. Access to retained experience differs from learning to use it.
Procedural memory migration Useful unchanged guidance matches paid selection at 12/12 with much lower total token cost. Recipient-dependent value is observed; harmful inheritance and paid adaptation repayment are not.
Correction lineage Selective repair touches fewer fields but costs more than rebuilding. Every arm repairs all states; later arithmetic failures leave complete behavior at 4/6.
Executable experience retention All four policies complete 8/8. Ready and archived source avoid generation; caching the first reconstruction removes later generation. The result concerns source availability, with acquired-lesson value and changed requirements untested.
Evidence use under revision One prespecified checkpoint transfers modestly and stays above base through corrections, with a favorable measured cost against an uncached lesson. Longer acquisition regresses; a supplied executable is complete.
Maintenance-decision transfer (assessment) Known history selects the two-support aggregate bound at 524/768 on fresh acquisitions, but every endpoint remains incomplete. Observation controls are constant; validation sees obligation differences despite aggregate ties. A retained-example table completes 192/192.
Experience selection — S1 (local) Damage-aware replay completes more uses overall than a developed fixed mixture, but fewer during recurrence, with more paired losses and substantial prediction cost. Both finish with complete acquisition; stopping at a fixed time can freeze incomplete learning.
Procedure retention and revision — S2 (local) A controlled crossing identifies an interaction between inherited state and current support, with the predicted direction in both fresh acquisitions. One support change helps one state and harms another; neither support choice fully preserves the second fresh acquisition.
Memory under goal shift — S4 (local) Better readers recover substantial utility for later compositions, while exact query-relevant omissions remain in some writers. Broad retention and explicit full records answer every tested query; losing some raw information need not prevent a particular later use.
Memory placement under revision — S5 (local) Learned reranking improves score predictions without repaying its cost: account-aware lexical retrieval matches full reranking at 127/128 complete answers, versus retained/invalidated EARM at 120/121. Prediction quality survives content edits better than changed evidence needs; complete useful advantage is absent.

S5 now includes an accepted placement comparison alongside the earlier acquisition, use and maintenance measurements. None demonstrates learning-cost repayment against competent explicit alternatives at comparable complete quality. The placement study shares visible evidence across methods; earlier references have different supplied structures and privileges. The later evidence-use result adds a favorable measured comparison against an uncached lesson, qualified above. These results do not establish a universal ranking of weights, records and executable rules.

Current work — 20 September 2026: Sixteen studies are registered. Fifteen have accepted bounded contributions and completed reviewed phases, including the lesson-acceptance pilot. The newly commissioned experience-guided investigation is prepared and pushed, ready for an independent session; no experiments have started. It tests whether an adapter can acquire investigation behavior from earlier attempts and use it on later unfamiliar changes after fresh session resets, with competent artifact and source reuse available. The task-sequence comparison selects ETL maintenance first and code search as a second lead. Reference checks pass; learner acquisition and added memory value remain untested. Software maintenance is a test setting within the enduring directive. The commission brief preserves CC1–CC3 and compares learning with ordinary work, including a useful alternative allocation of extra resources. Complete acquisition and transfer are primary; costs qualify their practical value. Root continues theory, literature and independent acceptance/intervention questions. The separate applied derivative owns prototype development; the lesson-acceptance pilot used its own SQLite/HTTP instrument and does not advance the runtime's review boundary. The next evidential milestone is useful capability across repeated realistic work and consequential change, with complete outcomes and total costs. The database follow-up establishes a narrower source-efficiency result, not that milestone. AD1 remains untested locally, AD2 has its bounded maintenance-choice assessment, and AD3 retains limited prediction-level support with adverse placement quality/cost evidence. EU, EX, PM, CL, WC and LA assessments supplement those predictions without replacing them.

S2's state/support phase is complete and accepted. MR1 has direct support in the selected equal-accuracy diagnostic; fresh acquisitions support the interaction but begin at different accuracies. MR2 and MR3 remain untested locally, with closer public precedents now identified. The reliability note preserves those expectations. S1, S2 and S4's completed phases remain stood down.

The learning-maintenance note preserves the original commissions and assesses their outcomes. The procedural-learning note records the new complete-preservation result without treating it as confirmation of reliable selective editability or repair of the earlier identifier-production problem. Failed acquisition deserves purposeful diagnosis; accepting a publication and closing its broader question remain separate decisions.

About

A continuation of agent memory research.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages