Lidor-Mashiach commited on
Commit
e3faafb
·
verified ·
1 Parent(s): 4a79c40

Add DeBERTa Large ANLI model

Browse files

Upload the fine tuned DeBERTa Large checkpoint, tokenizer, evaluation results, and model documentation.

Files changed (7) hide show
  1. README.md +164 -0
  2. baseline_eval.json +10 -0
  3. config.json +46 -0
  4. model.safetensors +3 -0
  5. model_card.json +35 -0
  6. tokenizer.json +0 -0
  7. tokenizer_config.json +18 -0
README.md CHANGED
@@ -1,3 +1,167 @@
1
  ---
2
  license: cc-by-nc-4.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: cc-by-nc-4.0
3
+ language:
4
+ - en
5
+ base_model: microsoft/deberta-large
6
+ datasets:
7
+ - facebook/anli
8
+ pipeline_tag: text-classification
9
+ tags:
10
+ - deberta
11
+ - anli
12
+ - natural-language-inference
13
+ - sequence-classification
14
+ - transformers
15
+ metrics:
16
+ - accuracy
17
+ model-index:
18
+ - name: DeBERTa Large ANLI
19
+ results:
20
+ - task:
21
+ type: text-classification
22
+ name: Natural Language Inference
23
+ dataset:
24
+ type: facebook/anli
25
+ name: ANLI combined rounds
26
+ split: dev
27
+ metrics:
28
+ - type: accuracy
29
+ value: 0.6221875
30
+ name: Validation accuracy
31
+ - task:
32
+ type: text-classification
33
+ name: Natural Language Inference
34
+ dataset:
35
+ type: facebook/anli
36
+ name: ANLI combined rounds
37
+ split: test
38
+ metrics:
39
+ - type: accuracy
40
+ value: 0.6178125
41
+ name: Test accuracy
42
  ---
43
+
44
+ # DeBERTa Large fine tuned on ANLI
45
+
46
+ ## Model
47
+
48
+ This checkpoint is based on `microsoft/deberta-large`.
49
+
50
+ It was fine tuned only on the combined ANLI training rounds. The training set contained 162,865 premise and hypothesis pairs.
51
+
52
+ The model predicts one of three labels:
53
+
54
+ | Label | Meaning |
55
+ |---:|---|
56
+ | 0 | entailment |
57
+ | 1 | neutral |
58
+ | 2 | contradiction |
59
+
60
+ The input order is premise first and hypothesis second.
61
+
62
+ ## Evaluation
63
+
64
+ Accuracy was measured on the combined ANLI held out rounds.
65
+
66
+ | Split | Accuracy | Examples |
67
+ |---|---:|---:|
68
+ | Development | 62.22% | 3,200 |
69
+ | Test | 61.78% | 3,200 |
70
+
71
+ These values are plain classification accuracy.
72
+
73
+ The checkpoint was trained on ANLI alone. Comparisons should use the same combined ANLI splits and the same label mapping. The results are not presented here as a universal leaderboard claim.
74
+
75
+ The machine readable results are stored in `baseline_eval.json`.
76
+
77
+ ## Training
78
+
79
+ | Setting | Value |
80
+ |---|---:|
81
+ | Base model | `microsoft/deberta-large` |
82
+ | Epochs | 3 |
83
+ | Batch size | 16 |
84
+ | Gradient accumulation steps | 2 |
85
+ | Learning rate | 0.000016958369168519958 |
86
+ | Weight decay | 0.1 |
87
+ | Warmup ratio | 0.1828387398995507 |
88
+ | Label smoothing | 0.1 |
89
+ | Maximum sequence length | 128 |
90
+ | Seed | 1299843651 |
91
+ | Numerical precision | FP32 |
92
+
93
+ The hyperparameters were selected for this model and dataset combination.
94
+
95
+ The full training record is stored in `model_card.json`.
96
+
97
+ ## Use
98
+
99
+ Load the repository with `AutoTokenizer` and `AutoModelForSequenceClassification` from the Transformers library.
100
+
101
+ Pass the premise and hypothesis as a text pair.
102
+
103
+ Use a maximum sequence length of 128 to match training.
104
+
105
+ ## Files
106
+
107
+ | File | Purpose |
108
+ |---|---|
109
+ | `model.safetensors` | Model weights |
110
+ | `config.json` | Architecture and label mapping |
111
+ | `tokenizer.json` | Tokenizer data |
112
+ | `tokenizer_config.json` | Tokenizer settings |
113
+ | `baseline_eval.json` | Evaluation results |
114
+ | `model_card.json` | Training record and provenance |
115
+ | `README.md` | Model card |
116
+
117
+ ## Limitations
118
+
119
+ The model was trained and evaluated on English ANLI data.
120
+
121
+ ANLI is adversarial and difficult. Performance on other NLI datasets may differ.
122
+
123
+ The training accuracy was 98.57%, while held out accuracy was lower. This gap should be considered when using the checkpoint.
124
+
125
+ The model can inherit errors and biases from the base model and the training data.
126
+
127
+ The checkpoint has not been evaluated for high risk or safety critical use.
128
+
129
+ ## License
130
+
131
+ The base model `microsoft/deberta-large` is licensed under MIT.
132
+
133
+ The ANLI training data is licensed under CC BY-NC 4.0.
134
+
135
+ This checkpoint is released under CC BY-NC 4.0 as a conservative noncommercial choice. Users must follow the terms of the base model and the ANLI dataset.
136
+
137
+ Use of this checkpoint is limited to noncommercial purposes.
138
+
139
+ ## Associated research
140
+
141
+ This model was trained as part of the following research manuscript:
142
+
143
+ **“Opening the Black Box: Localizing semantic inconsistency in NLI models with Deep k -Nearest Neighbors”**
144
+
145
+ The manuscript is in preparation. It has not been submitted or published.
146
+
147
+ This section will be updated when a public preprint or an accepted version becomes available.
148
+
149
+ ## Citation
150
+
151
+ Until the paper is public, please cite this model repository:
152
+
153
+ ```bibtex
154
+ @misc{mashiach2026debertaanli,
155
+ author = {Lidor Mashiach},
156
+ title = {DeBERTa Large fine tuned on ANLI},
157
+ year = {2026},
158
+ publisher = {Hugging Face},
159
+ url = {https://huggingface.co/Lidor-Mashiach/deberta-large-anli}
160
+ }
161
+ ```
162
+
163
+ Please also cite the DeBERTa and ANLI papers.
164
+
165
+ ## Contact
166
+
167
+ Questions, corrections, and reproducibility reports can be posted in the Community tab of this repository.
baseline_eval.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "DeBERTa-large",
3
+ "dataset": "ANLI",
4
+ "validation_accuracy": 0.6221875,
5
+ "validation_examples": 3200,
6
+ "validation_split": "dev",
7
+ "test_accuracy": 0.6178125,
8
+ "test_examples": 3200,
9
+ "test_split": "test"
10
+ }
config.json ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "DebertaForSequenceClassification"
4
+ ],
5
+ "attention_probs_dropout_prob": 0.1,
6
+ "bos_token_id": null,
7
+ "dtype": "float32",
8
+ "eos_token_id": null,
9
+ "hidden_act": "gelu",
10
+ "hidden_dropout_prob": 0.1,
11
+ "hidden_size": 1024,
12
+ "id2label": {
13
+ "0": "entailment",
14
+ "1": "neutral",
15
+ "2": "contradiction"
16
+ },
17
+ "initializer_range": 0.02,
18
+ "intermediate_size": 4096,
19
+ "label2id": {
20
+ "contradiction": 2,
21
+ "entailment": 0,
22
+ "neutral": 1
23
+ },
24
+ "layer_norm_eps": 1e-07,
25
+ "legacy": true,
26
+ "max_position_embeddings": 512,
27
+ "max_relative_positions": -1,
28
+ "model_type": "deberta",
29
+ "num_attention_heads": 16,
30
+ "num_hidden_layers": 24,
31
+ "pad_token_id": 0,
32
+ "pooler_dropout": 0.0,
33
+ "pooler_hidden_act": "gelu",
34
+ "pooler_hidden_size": 1024,
35
+ "pos_att_type": [
36
+ "c2p",
37
+ "p2c"
38
+ ],
39
+ "position_biased_input": false,
40
+ "relative_attention": true,
41
+ "tie_word_embeddings": true,
42
+ "transformers_version": "5.14.1",
43
+ "type_vocab_size": 0,
44
+ "use_cache": false,
45
+ "vocab_size": 50265
46
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fbcb3fc25a88581a85e9b9491e080abc8b4bd198f32c99dde5cc3bd1f10f8dd2
3
+ size 1624911148
model_card.json ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "DeBERTa-large",
3
+ "dataset": "ANLI",
4
+ "backbone": "microsoft/deberta-large",
5
+ "run_seed": 1299843651,
6
+ "validation_accuracy": 0.6221875,
7
+ "training": {
8
+ "epochs": 3,
9
+ "batch_size": 16,
10
+ "gradient_accumulation_steps": 2,
11
+ "learning_rate": 1.6958369168519958e-05,
12
+ "weight_decay": 0.1,
13
+ "warmup_ratio": 0.1828387398995507,
14
+ "label_smoothing_factor": 0.1,
15
+ "adam_beta2": 0.999,
16
+ "adam_epsilon": 1e-08,
17
+ "max_grad_norm": 1.0,
18
+ "max_seq_len": 128,
19
+ "fp16": "auto",
20
+ "amp": "fp32"
21
+ },
22
+ "baseline_eval": {
23
+ "validation_accuracy": 0.6221875,
24
+ "test_accuracy": 0.6178125,
25
+ "validation_examples": 3200,
26
+ "test_examples": 3200,
27
+ "validation_split": "dev",
28
+ "test_split": "test"
29
+ },
30
+ "repository": "Lidor-Mashiach/deberta-large-anli",
31
+ "license": "cc-by-nc-4.0",
32
+ "base_model_license": "mit",
33
+ "training_data": "facebook/anli",
34
+ "training_data_license": "cc-by-nc-4.0"
35
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": "[CLS]",
5
+ "cls_token": "[CLS]",
6
+ "do_lower_case": false,
7
+ "eos_token": "[SEP]",
8
+ "errors": "replace",
9
+ "is_local": true,
10
+ "local_files_only": false,
11
+ "mask_token": "[MASK]",
12
+ "model_max_length": 1000000000000000019884624838656,
13
+ "pad_token": "[PAD]",
14
+ "sep_token": "[SEP]",
15
+ "tokenizer_class": "DebertaTokenizer",
16
+ "unk_token": "[UNK]",
17
+ "vocab_type": "gpt2"
18
+ }