Ariadne Laya BioEvidence

Answers a biomedical research question with yes, no or maybe using the supplied abstract. This 4.2 MB specialist interface runs on a shared frozen Laya base. Switching between compatible Ariadne specialists replaces about 1 million parameters, while the 421 million parameter base stays in memory. Each specialist uses its own small interface.

Use

Install the included Python wheel from this downloaded model folder:

pip install ./ariadne_specialists-0.2.0a1-py3-none-any.whl
import ariadne
from ariadne.specialists import BioEvidence

model = ariadne.load_specialist(BioEvidence, model=".")
result = model({'question': 'Did the treatment reduce symptoms?', 'context': 'Symptoms were lower in the treatment group than in the placebo group.'})
print(result.label, result.score)

The base downloads automatically and is cached. Choose a device with device="cpu" or device="cuda:1". A list of inputs returns a list of results. Scores have not been recalibrated for this task.

Load from Hugging Face

After installing the included wheel, you can load this repository directly:

import ariadne
from ariadne.specialists import BioEvidence

model = ariadne.load_specialist(BioEvidence, model="GoatHerder/Ariadne-Laya-BioEvidence")

Use the explicit model= argument with this preview wheel. The interface and pinned base are downloaded automatically and cached. Pass revision="<commit hash>" to pin a particular interface version.

Interface

The interface is a 1,024 × 1,024 linear projection plus a 1,024-element bias: 1,049,600 trainable parameters. It sits after the base's native embeddings and before encoder block 0. It starts as the identity; training updates only this projection. The shared base has 421,293,827 parameters. Compatible specialists share one resident base in the same Python process and on the same device.

Results

Local evaluation uses the same 500 inputs for every model. Accuracy is per decision; macro-F1 averages the task's classes (and flag namespaces for Privacy).

Model Accuracy Macro-F1
Base Laya 50.20% 26.52%
Ariadne BioEvidence 61.40% 42.34%
nikhilteja30/pubmedqa-bert 93.00% 90.24%
TF-IDF + logistic regression 57.00% 33.88%

PubMedQA-named checkpoint with explicit yes/no/maybe labels. Its card does not document the dataset or pair formatting; provenance and input-format compatibility are unverified. Local comparison uses question/context pairs and is provisional. This table does not establish a common unseen-test ranking. metrics.json records model revisions, comparison methods, per-class results and existing-task retention.

Training and scope

Trained on qiaojin/PubMedQA, revision 9001f2853fb87cab8d220904e0de81ac6973b318. Prepared train/validation/test sizes: 450 / 50 / 500. Overlength exclusions: {'train': 0, 'validation': 0, 'test': 0}.

One seed (0); epoch 2 selected by validation loss, training stopped after epoch 10. LR 1e-4, minimum 10 epochs, patience 3. Only the 1,049,600 exact-identity-initialized interface parameters were trained. The original embeddings, 28 encoder blocks and decision heads stayed frozen and in evaluation mode. Deterministic GPU settings were enabled.

  • Official 500-example test PMID list; remaining 500 labeled examples split 450/50 with seed 42.
  • Input includes only question and abstract context; long_answer, final_decision and predicted reasoning labels are excluded. This is literature QA, not clinical advice.

English only. Unsupported or ambiguous inputs still receive a prediction. Source datasets retain their own licences. This checkpoint is one training run; it does not establish across-seed variance.

Base revision: 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Independent adaptation; no affiliation with the original Laya authors.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GoatHerder/Ariadne-Laya-BioEvidence

Adapter
(19)
this model

Dataset used to train GoatHerder/Ariadne-Laya-BioEvidence

Collection including GoatHerder/Ariadne-Laya-BioEvidence