FinBERT fine-tuned on financial tweets (RiskPulse)
Three-class financial sentiment (positive / negative / neutral) for short financial news posts and tweets. It is the sentiment component of RiskPulse, an individual hackathon project by Arnav Gupta (IIIT Dharwad).
Base model
ProsusAI/finbert, a BERT model further pre-trained on financial
text and fine-tuned for sentiment on Financial PhraseBank (Malo et al., 2014).
Training data
zeroshot/twitter-financial-news-sentiment (licence: MIT), English financial tweets labelled
Bearish / Bullish / Neutral, mapped to negative / positive / neutral.
Protocol (train split only)
- 8,597 rows (90% of the dataset's
trainsplit) for training; 946 rows (the other 10% oftrain) as a dev set for early stopping (best dev macro-F1 0.834 at epoch 2). - The dataset's
validationsplit (2,388 rows) is the test set and was evaluated once, after training. - Learning rate 2e-05, batch size 32, max length 64 tokens, 3 epochs, linear schedule with 10% warm-up, seed 20261002. CPU only, 45-minute budget (completed).
Results (test split, in-domain)
| Model | Macro-F1 |
|---|---|
| This model | 0.844 |
| Base FinBERT, labels from a score band tuned on the train split | 0.661 |
| Base FinBERT, argmax | 0.668 |
Leakage note. 52 of 2,388 test texts (2.2%) share their first 60 characters with a training text (39 are identical after removing URLs). That bounds the effect on macro-F1 at about 2 points, so it does not explain the gain.
In-domain caveat. Training and test texts come from the same dataset. Accuracy on other financial text (e.g. news headlines from other sources) has not been established by this number.
Intended use
Research and education: scoring the tone of short English financial headlines and posts as one input to a risk monitoring prototype. Not investment advice and not for automated trading decisions. Known limitations: short texts only (max 64 tokens in training), English only, the tone of a text is not its market impact.
Licence
Training data: MIT. The base model's card declares no licence; its code repository (ProsusAI/finBERT) is Apache-2.0, and its sentiment fine-tuning used Financial PhraseBank (CC BY-NC-SA 3.0). This derivative is therefore released under CC BY-NC-SA 3.0 as the conservative choice.
Labels
Model outputs positive, negative, neutral (see config.json → id2label). Dataset label ids were mapped as
{"0": "negative", "1": "positive", "2": "neutral"} (0 Bearish, 1 Bullish, 2 Neutral).
- Downloads last month
- 28
Model tree for arnavguptas/riskpulse-finbert-tweets
Base model
ProsusAI/finbert