Instructions to use Elda-AI/memory-resoner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Elda-AI/memory-resoner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="Elda-AI/memory-resoner", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Access to memory-resoner
Released for research, evaluation, internal validation and education. Tell us who you are and we will grant access.
By requesting access you agree to the Elda Community License 1.0: no commercial use, no redistribution of the weights or derivatives, and attribution as "Built with Elda". Commercial licensing is available on request.
Log in or Sign Up to review the conditions and access this model content.
memory-resoner β the reference step of conversational memory
A conversation can only be stored if you know what γκ·Έ λλ€γ meant. memory-resoner reads the turns of a conversation and says, for each mention, which earlier mention it refers to β so the layer above can write a memory entry that points at the right words.
"ν©μ λμ κ°κ²λ₯Ό μ΄μ΄μ . κ·Έ λλ€λ μ λμΈκ΅¬κ° λ§λμ ?"
mentions ν©μ λμ (s0, w0β0)
κ·Έ λλ€λ (s1, w0β1)
chains [ν©μ λμ Β· κ·Έ λλ€λ]
version 0.2.0 Β· 308M parameters Β· fp32 Β· 33 ms per sentence on CPU
when the mentions are supplied (125 ms when it finds them itself).
Where it sits
memory-resoner is one stage of a conversational stack, and it is multilingual β trained on Korean and English, measured on Japanese, and it transfers to languages it never saw. Each stage answers one question and hands the next stage a span, never prose.
utterance
β
ββ Elda-AI/intenter what was said, and what kind of thing it is ββ spans come from here
ββ slot extraction which of those the user said about themselves
β
βΌ
β
memory-resoner β
which earlier mention does this one refer to
β
βΌ
memory write a pointer into the user's own words β checkable, not generated
Where the spans come from. In the Elda stack, mention spans are produced upstream by Elda-AI/intenter and this model consumes them; its card is the reference for how a span is drawn, which types exist, and what the channels mean. Used standalone, memory-resoner will find its own mentions instead.
β Those two paths are not equally good, and the gap is large. On KLUE-WoS dialogue reference, this version scores 0.5074 when the mentions are supplied and 0.2547 when it has to find them itself. Detection and linking are two abilities and the end-to-end number is their product β so if you already have spans upstream, pass them in.
β The two sides count spans in different units, and the conversion is yours to make. Upstream spans are character offsets with the particle left outside (γν©μ λγ). This model works in μ΄μ (word) indices, and a μ΄μ contains its particle, so the same mention is γν©μ λμγ here. Neither is wrong; they are different units. Align on the μ΄μ that contains the upstream span.
What it returns. Span indices and chains. Not a rewritten sentence, not a summary β an index into the words you sent, which the layer above can verify before storing anything.
What it does not decide. Whether a fact is true, whether it is worth keeping, how long it lives. Those belong to the system around it. This stage answers one question and stops.
Why a small encoder
| this model | LLM prompting | |
|---|---|---|
| Parameters | 308M | 7B β 70B |
| Latency per sentence, CPU β linking, mentions given | p50 33 ms | seconds |
| Latency per sentence, CPU β finding mentions and linking | p50 125 ms | seconds |
| Output | span index β checkable byte-for-byte | free text to be parsed |
| Determinism | same input, same chains | sampling |
Memory is written on every turn. A stage that runs that often has to be cheap, and its answer has to be something the next stage can check rather than trust. A span index is both.
Output contract
input sentences, in order β a string per sentence, or a list of μ΄μ
output mentions Β· antecedents Β· chains (links closed transitively)
spans inclusive word (μ΄μ ) indices β Korean particles stay attached (see the note above)
window 6 previous sentences of context Β· 40 previous mentions as candidates
A mention with no antecedent opens a chain of its own β that is how the next stage learns it has seen a new entity, so singleton chains are kept rather than dropped.
Usage
from transformers import AutoModel, AutoTokenizer
m = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True).eval()
tok = AutoTokenizer.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True)
out = m.coref(["ν©μ λμ κ°κ²λ₯Ό μ΄μ΄μ .",
"κ·Έ λλ€λ μ λμΈκ΅¬κ° λ§λμ ?"], tok)
out["mentions"] # [{'sent': 0, 'words': [0, 0], 'text': 'ν©μ λμ'}, ...]
out["antecedents"] # [None, 0] β per mention: what it points back at
out["chains"] # [[0, 1]]
| call | does |
|---|---|
m.coref(sentences, tok) |
finds the mentions and links them |
m.mentions(sentences, tok) |
finds mentions only β a span provider for another linker |
m.resolve(units, mentions, tok) |
links only β you supply the mentions |
m.coref_last(sentences, tok) |
β serving shape β only the newest turn's mentions are returned, with the earlier turns kept as context |
m.resolve_last(sentences, mentions, tok) |
β serving shape, mentions supplied β the same, but you hand in the spans |
The two _last calls are what a live conversation wants: one turn arrives, you want answers for
that turn, and the turns before it are context rather than output. They take the whole
conversation and cut the window themselves, so the caller never has to know how long the window is.
β On
coref_last, do not count withantecedentsThe two calls differ here, and it matters.
resolve_lastreturns all the mentions you gave it, so itsantecedents[i]indexes that same list and reads normally.coref_lastreturns only the newest turn's mentions β so when the antecedent sits in an earlier turn, which on a live conversation is almost always,antecedents[i]isNone. There,Nonewears one face for two different facts: "it pointed at something outside this list" and "it pointed at nothing", and a consumer that counts non-Nonegets a structural zero even when the model is working.
field what it is use it for linked[i]β bool β did this mention resolve at all counting, on either call n_linkedsum of the above counting antecedent_mentions[i]the antecedent itself ( {sent, words, text}), orNonereading the answer antecedents[i]index into the returned list, or Nonesafe on resolve_last; oncoref_lastonly for within-turn links
resolve_lastalso reports what it had to drop or fix, and none of it is thrown away silently:mentions_outside_windowΒ·mentions_unencodableΒ·window_truncatedΒ·mentions_out_of_order. β Whenmentions_out_of_order > 0the layer re-sorted your mentions into reading order, and thenantecedents[i] > ican occur β the index is into the list you gave, not into reading order. If that count is zero, "an antecedent index is always smaller than its own" holds.
Notes that matter in practice:
- fp32. The config pins it. In bf16 near-ties flip and chains change.
- Pass sentences, not a paragraph. The sentence boundary is what the window is counted in.
- Send the resolved span downstream, not the reference. The output is an index into the words you sent, so whatever consumes it can verify what it was handed.
- Strip the particle at display time, not at span time. Keeping it inside the μ΄μ is what makes the span a plain index into your own input; trimming belongs to whatever renders the value.
Measurements
Public multilingual coreference, held out from training; document overlap with the training split is 0 and the benchmark script asserts it. Both scripts ship in this repository.
β Korean is one of the languages here, not the only one. If your pipeline is Korean-only, measure before you swap: reaching other languages costs some Korean headroom, and the trade is visible below. The two scripts here let you check that against your own data rather than take our word for it.
pip install torch transformers datasets
python bench_corefud_ko.py --model Elda-AI/memory-resoner
Mention detection β boundaries must match exactly
| precision | recall | F1 |
|---|---|---|
| 0.8354 | 0.6145 | 0.7081 |
Linking, mentions given
| scored mentions | 10263 |
| accuracy | 0.8305 |
| non-NULL accuracy | 0.4616 (n=2721) |
| deictic, non-NULL | 0.4207 (n=271) |
| reference β always NULL | 0.7349 |
| reference β always most-recent | 0.0448 |
Read the non-NULL rows. Most mentions open a chain, so answering NULL every time already scores 0.7349; the rows that require actually linking are the ones in bold.
End-to-end β (mention, antecedent) pairs, the model finding its own mentions
| precision | recall | F1 |
|---|---|---|
| 0.6554 | 0.508 | 0.5724 |
Every run first pushes the gold answers back through the scorer and asserts they survive intact (1.0000). A scorer that cannot return its own gold is not grading anything.
On dialogue
The numbers above are prose. On dialogue β references mined from a public multi-turn Korean corpus (KLUE-WoS), with the antecedent taken from that corpus's own annotation rather than chosen by us β this version resolves γκ·Έ μλΉγ-style references at 0.5074 when the mentions are supplied, and 0.2547 when it has to find them itself.
python bench_wos_references.py --model Elda-AI/memory-resoner
β That difference is the whole point of the two paths. End-to-end is detection times linking, so a pipeline that already has spans should pass them in rather than ask for them again. β These are measured with candidates mined by the gold rule, which makes them a perfect-upstream ceiling β a real intenter will score lower. These numbers are for conversation; on plain prose the same model reads reference less reliably.
What comes next
Conversational memory here is built in three layers. They are not future versions of this model β they are different kinds of component, and saying so is the point.
| what it is for | how it is built | |
|---|---|---|
| Layer 1 Β· conversation memory | what was said in this conversation, and what refers to what | β this model + deterministic assembly |
| Layer 2 | blackbox β how a memory is addressed and kept apart from every other conversation | not described here |
| Layer 3 Β· authority and persona | whether something holds and on what grounds, and what an agent is like across conversations | a deterministic VM, and a decoder used off the request path |
Layer 1 is the only one this model sits in. Layer 2 is deliberately not described: it is where memories are separated from one another, and it is the part of the system we keep closed. Layer 3 is where facts acquire grounds and where a persona is formed, and it is built from two very different machines β one that must be exact, and one that must be fluent.
The VM β the exact one. Memory is written in a small closed language and executed by a deterministic machine, so a write is either accepted, a duplicate, a replacement, or refused, and the same inputs always produce the same store. Reads are queries against that store, not recollection. This is what makes memory auditable rather than merely plausible, and it is why the per-turn path carries no generation at all.
The decoder β the fluent one. Never on the request path. Writing and reading memory happen on every turn and stay on the encoder-and-machine side. Generation is for prose a person will read, and for periodic passes that look across many conversations at once to form a persona β a pass that can fail without the conversation stopping.
What you can check yourself
| 0.2.0 | reproducible here? | |
|---|---|---|
| Japanese reference resolution β | 0.5664 | β |
| CorefUD, 19 languages, macro end-to-end | 0.1842 | β public data |
| in-domain span agreement (exact) β‘ | 97.5% | β |
| input positions this bundle uses | 384 | β config.max_len |
β JMultiWOZ 1.0 (CC BY-SA 4.0), references mined from that corpus's own slot annotation, words segmented with janome, 113 in-window cases, mentions supplied. β‘ measured on an in-house holdout of machine-generated dialogue, which we do not redistribute, so you cannot re-run it. It is listed because it is the axis our own consumers read, not as a claim you should take on trust. Rows marked β you can check yourself with the two scripts in this repository.
New knobs (all optional, all defaulting to the old behaviour)
| field | default | what it does |
|---|---|---|
detect_layers |
1 | BIO layers in the detection head. 1 emits outermost mentions only, as before. |
word_joiner |
" " |
What goes between words when the surface is rebuilt. Pass "" for languages written without spaces (Japanese, Chinese, Thai) β otherwise the encoder is handed a surface it has never seen. |
null_bias |
1.0 | How strongly to discourage the "no antecedent" choice. Larger means the model answers more often. Raise it only on a path where every mention really has an antecedent β see the warning below. |
| Japanese reference resolution | score |
|---|---|
word_joiner="" (as written) |
0.5664 |
word_joiner=" " (words spaced out) |
0.2301 |
Same weights, same inputs, same knob positions otherwise β far apart. A bigger encoder would not have saved you from this.
Each can also be passed per call: model.resolve_last(sentences, mentions, tok, joiner="", null_bias=2.0).
β Do not set
fix_mistral_regex=TrueRecent
transformersprints a warning when loading this tokenizer, suggesting the flag. We measured it on 5 sentences (ko, ja, en): with the flag, every one of them tokenizes to byte garbage βκ°λ¨μbecomesΓͺ Β° Δ· Γ« Δ€ Β¨instead ofβκ° λ¨ μ, and sequence length goes from 19 to 71. The model still runs and still returns spans; they are simply wrong. The default path is the correct one, and it is the one every number on this card was measured with.
β
null_bias: the number that looks free on the wrong test setOn a benchmark that only ever asks "which earlier mention does this pronoun refer to?", every abstention is wrong by construction, so pushing the model to always answer looks strictly better β and keeps looking better the harder you push. On text where "no antecedent" is sometimes the correct answer, the same push is destructive. We measured both:
null_biasKLUE-WoS, given mentions (no NULL-correct cases) CorefUD ko (NULL-correct cases present) 0 0.4409 0.5588 1 (shipped) 0.5074 0.5714 3 0.5443 0.5551 10 0.5616 0.1670
1.0is the only setting that improves both, and it does not trade away precision (accuracy among answered cases: 0.5576 β 0.5583). Going higher buys one number by destroying another. If your pipeline guarantees that the mentions you pass in are anaphoric, a higher value is safe; otherwise leave it alone.
License & access
Released under the Elda Community License 1.0 (see LICENSE). Provenance of the training data
is recorded in LINEAGE.md; no user conversations were used.
- β Research, evaluation, internal validation, education β free of charge
- β³ Attribution: "Built with Elda"
- β Commercial use and redistribution require a separate agreement
Access is gated: tell us who you are and access is granted automatically.
Citation
@software{memoryresoner2026,
title = {memory-resoner: multilingual reference resolution for conversational memory},
author = {Elda AI},
year = {2026},
url = {https://huggingface.co/Elda-AI/memory-resoner}
}
- Downloads last month
- 4