Goldenfish 1
Goldenfish 1 is a general-purpose decision model. It takes a state (a context plus a question) and a list of candidate answers, and returns a probability distribution over the candidates. It does not generate text. It reads, scores and picks.
The same model, with the same weights and the same call, covers decisions from two options to several hundred:
| Decision shape | Example question | Candidates |
|---|---|---|
| Yes / no | "Does the passage answer the question?" | yes, no |
| Few-option classification | "Which category applies?" | reset_password, billing, cancel_account, technical_issue |
| Intent detection | "What does the customer want?" | 10 to 150 intents |
| Priority | "How urgent is this ticket?" | low, medium, high |
| Policy | "Can this be published?" | allow, review, block |
| Tool selection | "Which tool should the agent call next?" | tool names or descriptions |
| Ticket / email / document classification | "Which queue handles this?" | queue names |
| Ranking | "Which reply is better?" | 2 to N responses |
| Routing | "Which agent handles this request?" | agent names |
| Large-label classification | "Which product category?" | hundreds of labels |
Candidates are plain strings. There is no training step to add a label: put it in the list.
- Parameters: 576,164,355 (backbone 566,705,152; decision heads 9,459,203)
- Backbone:
BAAI/bge-m3(XLM-RoBERTa large, 24 layers, 1024 hidden), revision5617a9f6 - Weights:
model.safetensors, float32, SHA256 inSHA256SUMS - Context window: up to 4,096 state tokens, 128 tokens per candidate
- Languages: trained on English, Portuguese, code and mixed text; the backbone is multilingual
Quick start
hf download aspaslabs/goldenfish-g1 --local-dir goldenfish-g1 # weights, tokenizer and code (about 2.3 GB)
pip install -e goldenfish-g1
Or install the package straight from the Hub (weights are fetched on first from_pretrained):
pip install -U pip
pip install "git+https://huggingface.co/aspaslabs/goldenfish-g1"
Requires Python 3.10+ and pip 23+ (older pip versions can fail to resolve the torch dependency tree). Verified on Python 3.10 and 3.12 with transformers 4.4x and 5.x (CPU); CUDA inference uses the same code path via device="cuda".
Goldenfish.from_pretrained accepts either the local directory or the repo id; with the repo id, the files are fetched from the Hub and cached.
from goldenfish import Goldenfish
model = Goldenfish.from_pretrained("aspaslabs/goldenfish-g1") # device="cuda", dtype=torch.bfloat16 optional
decision = model.decide(
context="forgot my password",
question="Which category applies?",
candidates=["reset_password", "billing", "cancel_account", "technical_issue"],
)
print(decision.choice, decision.confidence)
# reset_password 0.95
decide returns a Decision:
decision.candidates # the list you passed, same order
decision.probabilities # softmax over candidates, sums to 1
decision.index # argmax
decision.choice # candidates[index]
decision.confidence # probabilities[index]
decision.margin # top-1 minus top-2 probability
decision.entropy # nats
decision.ranked(k=3) # [(candidate, probability), ...]
More examples:
# Banking dispute
model.decide("this charge wasn't mine", "Which category applies?",
["fraud", "refund", "card_delivery", "cash_withdrawal"]).choice
# fraud
# Policy decision
model.decide("Draft quotes a customer's full name and home address without consent.",
"Can this be published?", ["allow", "review", "block"]).choice
# Yes / no, no context
model.decide("", "Is Lisbon the capital of Portugal?", ["yes", "no"]).choice
# Batch
model.decide_batch([
("The API returns 502 for every request since the deploy.", "What is the priority?", ["low", "medium", "high"]),
("Can you change my delivery address?", "Which team?", ["shipping", "billing", "returns"]),
])
Static candidate pools
Candidate embeddings do not depend on the state. For a fixed label set, encode it once and reuse it; subsequent calls run the encoder only on the state.
labels = load_my_300_intents()
model.cache_candidates(labels)
model.rank("I need a flight to Rome next Tuesday", "Which intent?", labels, k=3)
With a cached pool, a 301-candidate decision costs the same forward pass as a 2-candidate one, plus a 301 x 512 matrix product and a small refinement over the top 8.
examples.py in this repository runs all of the above.
Architecture
Goldenfish is a bi-encoder with a set-level refinement head.
state = question [SEP] context -> encoder -> CLS -> MLP -> LayerNorm -> L2 norm s (512-d)
cand_i = candidate text -> encoder -> CLS -> MLP -> LayerNorm -> L2 norm c_i (512-d)
base_i = <s, c_i> / tau tau learned, 0.045 in this release
logits = base + SetMixer(s, {c_i}, base)[top-8 of base]
p = softmax(logits)
- Encoder: the two sides share one
BAAI/bge-m3encoder; state and candidate have separate projection heads. - Scoring: cosine similarity scaled by a learned temperature. This is the separable part: candidates can be precomputed, and scoring N candidates is one matrix product.
- Set mixer: a 2-layer Transformer encoder over the set {state, top-8 candidates}. Each candidate token carries its embedding plus its first-stage score. There are no positional embeddings, so the refinement is equivariant to candidate order. Its output is a residual added to the top-8 logits. With more than 8 candidates the rest keep their first-stage scores; recall@8 of the first stage on our high-cardinality development sets is above 0.988, so the cut-off is not a bottleneck.
- Output: softmax over the candidates present. Padding positions are masked.
Everything is in goldenfish/modeling.py (about 200 lines) and goldenfish/inference.py. No custom CUDA, no remote code.
Goldenfish and Julia-1
Julia-1 (SupersonicLabs) is the closest published model with the same input contract (state, question, options) and the same no-generation output. The two make a different architectural trade.
Julia-1 is a joint encoder: the state and every option are encoded together in one sequence, so every attention layer sees both. That gives full state-option interaction and is very effective when there are a few options whose meaning depends on the state (board positions, "will this move collide", "does this hypothesis follow from this premise"). The cost grows with the number of options because the state is re-read for each one, and options cannot be precomputed.
Goldenfish encodes state and candidates separately and lets them interact only through the dot product and the set mixer. That makes candidate scoring separable and cacheable: one state forward pass serves any number of candidates, a fixed label set is encoded once, and latency is flat in the number of candidates. It is the natural shape for classification, intent detection, routing, ranking and candidate scoring over large pools, and it works at two candidates as well: yes/no, allow/review/block and low/medium/high are first-class uses, not edge cases.
The honest reading of the benchmarks below is that this trade is visible in the numbers. Goldenfish is ahead on the typed-decision suite overall and on the public classification sets, and it is behind Julia-1 on Open-Jev Snake, a low-cardinality task that needs the state to be read against each option. Improving low-cardinality state-candidate reasoning without giving up separable scoring is the main line of work for the next release.
Benchmarks
All numbers below are for this exact checkpoint. Goldenfish numbers were produced by us; Julia-1 numbers are from the Julia-1 model card unless marked otherwise. Selection of this checkpoint used only training-derived development sets; the test sets below were opened once, after the checkpoint was frozen.
Typed decisions (Julia-1 protocol, 2,000 test rows)
| Goldenfish 1 | Julia-1 | |
|---|---|---|
| Overall accuracy | 73.20 | 73.15 |
| Choice | 76.75 | 80.60 |
| Noul (yes/no) | 76.20 | 73.80 |
| Score | 66.65 | 69.10 |
The overall margin is one example in 2,000; treat it as parity. Goldenfish is ahead on Noul and behind on Choice and Score.
Public classification sets (100-example pilots, Jev protocol)
| Goldenfish 1 | Julia-1 | |
|---|---|---|
| AG News (4 classes) | 95 | 89 |
| DAIR Emotion (6 classes) | 90 | 85 |
| Banking77, 72-way, 100 examples | 90 | not reported in this form |
These are 100-example pilots and carry roughly plus or minus 5 points of sampling noise. Banking77 is listed for information only; Julia-1 reports a different Banking77 protocol.
MASSIVE intent (60 intents)
| Goldenfish 1 | Julia-1 | |
|---|---|---|
| Global | 82.70 | 80.55 |
| pt-PT | 87.29 | 84.52 |
| en-US | 88.03 | 85.57 |
Internal development suite
Full development set, 14,802 rows, 30 datasets across 7 families (entailment, contextual yes/no, intent routing, ordinal scoring, policy and planning, ranking and preference, commonsense multiple choice): 81.00% accuracy, NLL 0.48, ECE 0.012. Strong families: policy and planning 97.3, intent routing 90.7, contextual yes/no 83.2, ordinal 82.0. Weak families: commonsense multiple choice (ARC, HellaSwag, CommonsenseQA, OpenBookQA) 47.2, preference ranking 66.3. By number of candidates: K<=2 81.8, K<=3 84.8, K<=16 90.7, K<=32 85.5, K<=64 78.2, K>77 72.2.
Open-Jev Snake (public validation, 848 rows)
ZefanCai/Open-Jev, config release-v2-redistributable, source snake-v1, validation split, revision c67699e1. Both models run zero-shot on the same rows, same A100, batch size 1.
| Goldenfish 1 | Julia-1 | |
|---|---|---|
| Total (785 one-hot rows) | 51.85 | 86.11 |
| Choice (149) | 40.27 | 82.55 |
| Noul (636) | 54.56 | 86.95 |
| NLL | 0.796 | 0.412 |
| p50 latency, end to end | 13.5 ms | 10.3 ms |
| Peak VRAM | 1.17 GB | 0.79 GB |
This is a clear loss. The task asks whether a move on a 6x6 board collides, from a JSON board state; it needs the option to be read against the geometry of the state, which is exactly what a separable scorer does not do. Julia-1's published figure for this source (79.0% on an unpublished 200-row subset) is not directly comparable to the 848-row public run. Snake-v1 appears in Julia-1's public development and evaluation provenance; Goldenfish was not trained on it.
Latency
Single decision, batch size 1, A100, float32 weights with bf16 autocast, 4 candidates of up to 128 tokens and a state of about 100 tokens: p50 13.5 ms end to end, 73 decisions per second, 1.17 GB peak VRAM. On an 8-core CPU in float32: about 0.2 s per decision with 4 fresh candidates, and about 80 ms for a 301-way decision over a cached label pool, where the state is the only text encoded.
Intended use
Goldenfish is meant to sit where a system must choose among known options and wants a calibrated probability for each:
- Support and operations: ticket triage, queue routing, priority assignment, intent detection over a live intent catalogue.
- Agents: picking the next tool, the next agent, or the next step from a declared set; policy gates such as allow/review/block before an action.
- Content and compliance: document and email classification, publishability checks, policy categories.
- Ranking: choosing the better of N candidate responses or documents for a given request.
- Large label spaces: product taxonomies, hundreds of intents or FAQ entries, with labels encoded once and cached.
Candidate text matters. Descriptive candidates (refund for a duplicate charge) are scored more reliably than opaque codes (CAT_17). Questions should be phrased as the decision being made, in natural language, and kept constant for a given use case.
Limitations
- Low-cardinality state-option reasoning that requires reading the state against each option (geometry, numeric comparison, strict premise-hypothesis logic) is the weakest area. Open-Jev Snake above is the clearest case. If your task is of that kind and has few options, a joint encoder such as Julia-1 is likely to do better today.
- Knowledge-heavy multiple choice (ARC, HellaSwag, CommonsenseQA) is close to 40 to 50%. Goldenfish is not a knowledge model and does not reason over long chains.
- Confidence is calibrated on our development suite (ECE 0.012 overall) but can be high on wrong answers for entailment-style questions; use
marginas well asconfidencefor abstention. - Candidate truncation at 128 tokens and state truncation at 4,096 tokens are silent.
- The 100-example public pilots are indicative, not definitive.
- The model was trained on English, Portuguese and code. Other languages inherit whatever the backbone provides and have not been evaluated.
Files
README.md this model card
LICENSE Apache-2.0
CITATION.cff
SHA256SUMS checksums of the weights and tokenizer files
config.json Goldenfish configuration (pooling, projection, mixer, temperatures, limits)
backbone_config.json XLM-RoBERTa configuration of the bge-m3 encoder (pinned, no Hub access needed to build the model)
model.safetensors 576,164,355 float32 parameters
tokenizer.json, tokenizer_config.json, special_tokens_map.json, sentencepiece.bpe.model
goldenfish/ inference package (modeling.py, inference.py)
pyproject.toml pip-installable package definition
examples.py runnable examples
Dependencies: torch>=2.2, transformers>=4.40, safetensors, sentencepiece, huggingface_hub.
Training summary
Goldenfish 1 was trained on a mixture of public decision-shaped datasets reformatted as (context, question, candidates, gold) rows: natural language inference (MNLI, SNLI, ASSIN2), yes/no QA (BoolQ, ExtraGLUE BoolQ-pt), intent and scenario classification (Banking77, CLINC150, MASSIVE), preference (HH-RLHF), ordinal similarity (STS-B, ASSIN2), commonsense multiple choice, a typed-decision corpus following the Julia-1 protocol, and synthetic agentic routing, tool-selection and policy rows. Only training splits were used for fitting; development sets were derived from training data. The objective combines cross-entropy over candidates, KL to soft targets where available, and ranking and consistency terms. The released weights are a 0.95/0.05 interpolation between a final supervised checkpoint and an earlier external-transfer checkpoint, chosen on training-derived development data before any test set was opened.
License
The weights and code in this repository are released under the Apache License 2.0. The backbone BAAI/bge-m3 is released by BAAI under the MIT license.
Citation
@misc{goldenfish1,
title = {Goldenfish 1: a general-purpose decision model},
author = {Basecase Labs},
year = {2026},
url = {https://huggingface.co/aspaslabs/goldenfish-g1}
}
- Downloads last month
- -
Model tree for aspaslabs/goldenfish-g1
Base model
BAAI/bge-m3