FinBERT fine-tuned on financial tweets (RiskPulse)

Three-class financial sentiment (positive / negative / neutral) for short financial news posts and tweets. It is the sentiment component of RiskPulse, an individual hackathon project by Arnav Gupta (IIIT Dharwad).

Base model

ProsusAI/finbert, a BERT model further pre-trained on financial text and fine-tuned for sentiment on Financial PhraseBank (Malo et al., 2014).

Training data

zeroshot/twitter-financial-news-sentiment (licence: MIT), English financial tweets labelled Bearish / Bullish / Neutral, mapped to negative / positive / neutral.

Protocol (train split only)

  • 8,597 rows (90% of the dataset's train split) for training; 946 rows (the other 10% of train) as a dev set for early stopping (best dev macro-F1 0.834 at epoch 2).
  • The dataset's validation split (2,388 rows) is the test set and was evaluated once, after training.
  • Learning rate 2e-05, batch size 32, max length 64 tokens, 3 epochs, linear schedule with 10% warm-up, seed 20261002. CPU only, 45-minute budget (completed).

Results (test split, in-domain)

Model Macro-F1
This model 0.844
Base FinBERT, labels from a score band tuned on the train split 0.661
Base FinBERT, argmax 0.668

Leakage note. 52 of 2,388 test texts (2.2%) share their first 60 characters with a training text (39 are identical after removing URLs). That bounds the effect on macro-F1 at about 2 points, so it does not explain the gain.

In-domain caveat. Training and test texts come from the same dataset. Accuracy on other financial text (e.g. news headlines from other sources) has not been established by this number.

Intended use

Research and education: scoring the tone of short English financial headlines and posts as one input to a risk monitoring prototype. Not investment advice and not for automated trading decisions. Known limitations: short texts only (max 64 tokens in training), English only, the tone of a text is not its market impact.

Licence

Training data: MIT. The base model's card declares no licence; its code repository (ProsusAI/finBERT) is Apache-2.0, and its sentiment fine-tuning used Financial PhraseBank (CC BY-NC-SA 3.0). This derivative is therefore released under CC BY-NC-SA 3.0 as the conservative choice.

Labels

Model outputs positive, negative, neutral (see config.json → id2label). Dataset label ids were mapped as {"0": "negative", "1": "positive", "2": "neutral"} (0 Bearish, 1 Bullish, 2 Neutral).

Downloads last month
28
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arnavguptas/riskpulse-finbert-tweets

Base model

ProsusAI/finbert
Finetuned
(113)
this model

Dataset used to train arnavguptas/riskpulse-finbert-tweets