🇧🇩 BanglaTuring: Bengali AI-Generated Text Detector (BanglaBERT-SupCon)

Hugging Face Model Base Model License Task

A production-grade, generator-resilient sequence classification model for detecting AI-generated Bengali text.

Fine-tuned on BanglaBERT (csebuetnlp/banglabert) using Supervised Contrastive Learning (SupCon) and calibrated with Temperature Scaling ($T = 1.8816$), this model accurately separates machine-generated Bengali text from authentic human writing across modern frontier LLMs.


🌟 Why Use This Model?

Generic multilingual models (such as mBERT or XLM-RoBERTa) often break Bengali text into arbitrary 3 to 5 subword fragments and overfit to superficial generator formatting templates. This model is engineered specifically for the nuances of Bengali:

  1. Native Bengali Pretraining: Built upon BUET's BanglaBERT (ELECTRA architecture) with a dedicated 32,000-token Bengali vocabulary, preserving complete morphological roots and grammatical inflections.
  2. Trained Across Modern Frontier Generators:
    Universal AI Manifold (SupCon)
    ├── ChatGPT (GPT-5.6 Luna)
    ├── Claude (Sonnet 4)
    ├── DeepSeek (DeepSeek-V4)
    ├── Google (Gemini 3.1 Pro)
    └── Grok (Grok 4.1)
    
  3. Resilient to Generator Shift (SupCon): Rather than memorizing the signature introductory phrases or paragraph habits of a single LLM, Supervised Contrastive Learning pulls synthetic Bengali representations into a unified hypersphere while repelling human prose. It retains high detection sensitivity even on newly released or unseen AI models.
  4. Calibrated Probabilities (Zero False Overconfidence): Raw deep learning models often output false 99.9% certainty. Using temperature scaling ($T = 1.8816$), confidence scores reflect true empirical likelihoods, safeguarding innocent human writers against false accusations.

📊 Core Performance Highlights

Evaluated on rigorous Leave-One-Generator-Out (LOGO) cross-generator validation and verified on independent multi-generator benchmarks:

  • In-Distribution Accuracy: $98.50%$
  • Cross-Generator Unseen AI Recall: $95.04%$ (consistently catching AI text even from models withheld entirely from training)
  • Human Specificity: $95.99%$ (minimizing false alarms on formal, academic, and creative human Bengali writing)
  • External Zero-Shot Transfer: $98.03%$ accuracy across independent external holdouts

🚀 Quick Start & Usage

1. Simple High-Level Pipeline (3 Lines)

from transformers import pipeline

# Initialize classifier
classifier = pipeline(
    "text-classification",
    model="MahatirTusher/bangla-ai-text-detector",
    return_all_scores=True
)

# Test sample
text = "কৃত্রিম বুদ্ধিমত্তা বর্তমান যুগে তথ্যপ্রযুক্তির এক অভাবনীয় বিপ্লব ঘটিয়েছে।"
result = classifier(text)

print(result)
# Output: [[{'label': 'Human', 'score': 0.018}, {'label': 'AI', 'score': 0.982}]]

2. PyTorch Inference with Temperature Calibration (Recommended)

For production workflows, applying temperature scaling ($T = 1.8816$) produces calibrated, well-balanced probabilities:

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

MODEL_NAME = "MahatirTusher/bangla-ai-text-detector"
TEMPERATURE = 1.8816  # Calibrated scaling factor

tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_NAME)
model.eval()

def detect_bengali_text(text: str, threshold: float = 0.50):
    inputs = tokenizer(
        text,
        return_tensors="pt",
        truncation=True,
        max_length=256
    )
    
    with torch.no_grad():
        logits = model(**inputs).logits[0]
        # Temperature scaling
        calibrated_logits = logits / TEMPERATURE
        probs = torch.softmax(calibrated_logits, dim=-1)
        
    human_prob = float(probs[0].item())
    ai_prob = float(probs[1].item())
    
    verdict = "AI-generated" if ai_prob >= threshold else "Human-written"
    
    # Assess certainty based on margin from decision boundary
    margin = abs(ai_prob - threshold)
    confidence = "Very High" if margin >= 0.35 else "High" if margin >= 0.20 else "Moderate"
    
    return {
        "verdict": verdict,
        "ai_probability": round(ai_prob, 4),
        "human_probability": round(human_prob, 4),
        "confidence": confidence
    }

# Example
sample = "বাংলাদেশ দক্ষিণ এশিয়ার একটি নদীমাতৃক ও সার্বভৌম রাষ্ট্র।"
print(detect_bengali_text(sample))

🎯 Operating Modes & Recommended Thresholds

Depending on your application, you can adjust the decision threshold ($\tau$):

Operating Mode Decision Threshold ($\tau$) Best Suited For
Balanced (Default) 0.50 General web text, blogs, social media posts, news analysis.
High Precision 0.75 – 0.90 Academic integrity, publishing, and legal checks (strictly minimizes false alarms).
High Recall 0.35 Automated spam screening and preliminary community moderation.

💡 Best Practices for Long Documents

  • Short Texts ($< 25$ words): Short fragments carry fewer stylistic and contextual markers. Provide complete sentences for optimal reliability.
  • Long Documents ($> 120$ words): For long articles, essays, or reports, consider a sentence-snapped sliding window inspection ($100–120$ words per window with $30–40$ words overlap). This allows you to pinpoint specific synthetic sections in partially AI-assisted documents.

⚠️ Probabilistic Nature & Responsible Use

No automated detection system can determine authorship with 100% certainty. This model outputs statistical probabilities based on patterns learned across large corpora. Results should always be interpreted in conjunction with human judgment, contextual awareness, and domain knowledge.


👨‍💻 Authors & Attribution

  • Principal Investigator & Model Developer: Mahatir Ahmed Tusher
  • AI Data Generation & Curation: Sagar Chandra Dey
  • Initiative: Khoj — Advanced AI Fact-Checking & Digital Information Integrity
  • Base Model: BUET BanglaBERT
@misc{tusher2025bengaliaidetector,
  author = {Mahatir Ahmed Tusher and Sagar Chandra Dey},
  title = {BanglaTuring: Bengali AI-Generated Text Detector via Supervised Contrastive Learning},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/MahatirTusher/bangla-ai-text-detector}}
}

📄 License

This model and its associated inference code are distributed under the MIT License.

Downloads last month
102
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support