OZZZER · AI NEWS1 of 3 free stories opened
← Back to AI News

Coding & Development · 3 Sep 2026 · 02:00 CEST

Give Your Coding Agents a Memory You Own

Hugging Face · 3 Sep 2026 · 02:00 CESTRead original at Hugging Face ↗
Share
LinkedInX
Give Your Coding Agents a Memory You Own

Publisher preview · OZZZER analysis pending editorial review.

PUBLISHER ARTICLE PREVIEW

From the original article

Earlier this year, Software Forgets: Agent Traces Are the Memory made the case that coding agents already produce the record we keep losing. As they search a codebase, try approaches, hit errors, read documentation, and change direction, they leave behind a dense account of not just what changed, but why.

While the diagnosis is correct, traces are only potential memory. The session logs of an agent are still just an archive. You cannot grep your way to “why did we move off the streaming parser?” across ten thousand turns. For an agent to use those traces while it works, they need indexing, retrieval, ranking, and exact provenance.

That is what funes provides. It is a durable memory layer for your agents (Claude Code, Codex, pi, and Hermes). It is built from the sessions already on your machine. It works locally and becomes part of your agent's normal workflow with one command. When you want it to, it can also travel to a Hugging Face dataset you own, private by default.

funes is a single binary. Its default inference backend has no ML runtime dependency, and embedding and reranking happen on your machine. Install it:

That one add command builds the first index, gives the agent recall and get tools, and installs the automation that indexes each completed turn. Indexing is incremental, with new runs adding new turns rather than embedding the whole history again. The older and deeper content can backfill in bounded steps.

From there, you just work. When a task touches a past decision, rationale, or finding, the agent can reach for recall itself. You do not need to remember the old session or paste its context into the new one.

With funes added, recall happens inside the conversation. The agent reaches for its memory on its own and names the session behind its answer.

recall returns the original text, not a summary, and shows exactly where it came from (the agent, timestamp, session, and turn). Each result includes a get command that opens the full turn and its surrounding context.

Underneath, one deterministic pipeline parses every supported trace into the same turn-and-block shape, chunks it, embeds it with a pinned local model, and writes it to a local Lance dataset. A query combines vector and BM25 search, fuses their rankings, reranks the candidates with a cross-encoder, reweights them by recency, and attaches neighboring chunks.

The agent as a stranger problem is already solved on one machine. But memory gets more useful when the next agent is running somewhere else.

To make a memory follow your work, bind one when you

Source

Hugging Face · 3 Sep 2026 · 02:00 CEST

Open the original at Hugging Face ↗