END-TO-END APPLIED NLP SYSTEM
A high-precision sentiment pipeline fine-tuned on IMDB, evaluated against an uncompromising TF-IDF baseline, and deployed with a low-latency FastAPI inference server.
SYSTEM ARCHITECTURE
Three methodical stages from raw dataset exploration to high-throughput serving.
Trained on 22,500 reviews with 10k n-gram features. Achieves 89.20% accuracy in 0.38 ms — setting an uncompromising baseline that proves whether deep learning is justified.
Fine-tuned with HuggingFace Trainer, PyTorch, and Weights & Biases tracking. Configured with max_seq_length=256 to capture full semantic context without review truncation.
Asynchronous REST service featuring in-memory model registry, dynamic tensor padding, and true batched inference delivering a 5.9× throughput improvement.
EMPIRICAL VALIDATION
Evaluated on the exact same 25,000-example held-out test split. No simulated numbers or hypothetical projections.
| Evaluation Metric | TF-IDF Baseline | DistilBERT (Fine-Tuned) | Delta & Impact |
|---|---|---|---|
| Accuracy | 89.20% | 91.30% | +2.10 pp |
| Precision | 88.81% | 90.98% | +2.16 pp |
| Recall | 89.69% | 91.70% | +2.01 pp |
| F1-Score | 89.25% | 91.33% | +2.09 pp |
| Single Latency | 0.38 ms (CPU) | 15.0 ms (MPS) | 41× slower |
| Batch Latency (B=32) | 0.38 ms / item | 1.70 ms / item | 5.9× speedup |
| Disk Footprint | 0.46 MB | 267.8 MB | 585× larger |
| Recommended Use | Edge / <1ms SLA | Accuracy-First SLAs | Contextual |
The TF-IDF baseline achieves 89.20% accuracy in sub-millisecond latency (0.38 ms) with negligible memory footprint (0.46 MB). Ideal for high-concurrency microservices (>5k req/sec) on modest CPU hardware.
DistilBERT provides an essential +2.10 pp accuracy boost on complex, nuanced sentiments. Combined with batching (5.9× speedup at 1.7 ms/item), it provides exceptional throughput for offline pipelines and high-value decisioning.
PRACTICAL TRADE-OFFS
Real architectural takeaways, root-cause debugging narratives, and production decisions documented during development.
In applied NLP, a well-tuned TF-IDF baseline is notoriously competitive (89.20% accuracy at 0.38 ms). Fine-tuning transformer models is only justifiable if it statistically clears that threshold. ForgeML measures whether +2.10 pp accuracy is worth a 41× single-sample latency cost.
Our initial fine-tuning run scored 87.83%—worse than the baseline. Root-cause analysis revealed IMDB reviews average 230 tokens; a 128-token limit discarded vital sentiment shifts in final acts. Expanding context to 256 tokens unlocked 91.30%, proving data distributions dictate hyperparameters.
Instead of jumping to multi-billion parameter LLMs with massive latency overhead, we optimized inference throughput through true tensor batching. By passing attention masks and executing unified matmuls, batch-of-32 achieves a 5.9× per-item speedup (1.70 ms/item).
The FastAPI serving layer isolates model lifecycles behind an in-memory registry singleton (get_model). This decouples request routing from tensor loading, guarantees zero per-request initialization overhead, and allows seamless fallback between local weights and HuggingFace Hub.
Calls POST /predict on the live FastAPI backend. No simulated latency or cached results.