ForgeML

END-TO-END APPLIED NLP SYSTEM

SENTIMENT,UNDER-STOOD.

A high-precision sentiment pipeline fine-tuned on IMDB, evaluated against an uncompromising TF-IDF baseline, and deployed with a low-latency FastAPI inference server.

++++
TELEMETRY / BENCHMARKVERIFIED
TEST ACCURACY91.30%+2.10 pp vs Baseline
MPS LATENCY15.0msSingle Prediction
Model Arch:distilbert-base-uncased
Max Context:256 tokens
Batch Speedup:5.9× (B=32 @ 1.7ms/item)
CORE TECHNOLOGY STACKPRODUCTION INFRASTRUCTURE
Deep LearningPyTorch 2.13
Tokenization & TrainerHuggingFace Transformers
Experiment TelemetryWeights & Biases
Async REST ServingFastAPI & Uvicorn
App Router FrontendNext.js 16 (React 19)
Procedural CanvasThree.js / R3F
Motion OrchestrationGSAP & ScrollTrigger
Design System TokensTailwind CSS v4

SYSTEM ARCHITECTURE

How It Works

Three methodical stages from raw dataset exploration to high-throughput serving.

01
TF-IDF + LogisticRegression

Classical Baseline

Trained on 22,500 reviews with 10k n-gram features. Achieves 89.20% accuracy in 0.38 ms — setting an uncompromising baseline that proves whether deep learning is justified.

02
DistilBERT Sequence Classifier

Transformer Fine-Tuning

Fine-tuned with HuggingFace Trainer, PyTorch, and Weights & Biases tracking. Configured with max_seq_length=256 to capture full semantic context without review truncation.

03
FastAPI + Batch Engine

Production Serving

Asynchronous REST service featuring in-memory model registry, dynamic tensor padding, and true batched inference delivering a 5.9× throughput improvement.

EMPIRICAL VALIDATION

Model Comparison

Evaluated on the exact same 25,000-example held-out test split. No simulated numbers or hypothetical projections.

Evaluation MetricTF-IDF BaselineDistilBERT (Fine-Tuned)Delta & Impact
Accuracy89.20%91.30%+2.10 pp
Precision88.81%90.98%+2.16 pp
Recall89.69%91.70%+2.01 pp
F1-Score89.25%91.33%+2.09 pp
Single Latency0.38 ms (CPU)15.0 ms (MPS)41× slower
Batch Latency (B=32)0.38 ms / item1.70 ms / item5.9× speedup
Disk Footprint0.46 MB267.8 MB585× larger
Recommended UseEdge / <1ms SLAAccuracy-First SLAsContextual
WHEN TO USE BASELINE

High-Throughput & Low-Compute Edge

The TF-IDF baseline achieves 89.20% accuracy in sub-millisecond latency (0.38 ms) with negligible memory footprint (0.46 MB). Ideal for high-concurrency microservices (>5k req/sec) on modest CPU hardware.

Footprint: 0.46 MB · Latency: 0.38 ms
WHEN TO USE DISTILBERT

Accuracy-Critical & Batch Scoring

DistilBERT provides an essential +2.10 pp accuracy boost on complex, nuanced sentiments. Combined with batching (5.9× speedup at 1.7 ms/item), it provides exceptional throughput for offline pipelines and high-value decisioning.

Accuracy: 91.30% · Batching: 5.9× Speedup

PRACTICAL TRADE-OFFS

Engineering Decisions

Real architectural takeaways, root-cause debugging narratives, and production decisions documented during development.

[ARCHITECTURE]FORGEML // LOG

Why Compare Against a Baseline First

In applied NLP, a well-tuned TF-IDF baseline is notoriously competitive (89.20% accuracy at 0.38 ms). Fine-tuning transformer models is only justifiable if it statistically clears that threshold. ForgeML measures whether +2.10 pp accuracy is worth a 41× single-sample latency cost.

Standard VerifiedSection Ref ✓
[DIAGNOSIS]FORGEML // LOG

The max_seq_length Truncation Bug

Our initial fine-tuning run scored 87.83%—worse than the baseline. Root-cause analysis revealed IMDB reviews average 230 tokens; a 128-token limit discarded vital sentiment shifts in final acts. Expanding context to 256 tokens unlocked 91.30%, proving data distributions dictate hyperparameters.

Standard VerifiedSection Ref ✓
[PERFORMANCE]FORGEML // LOG

Batching Over Bigger Models

Instead of jumping to multi-billion parameter LLMs with massive latency overhead, we optimized inference throughput through true tensor batching. By passing attention masks and executing unified matmuls, batch-of-32 achieves a 5.9× per-item speedup (1.70 ms/item).

Standard VerifiedSection Ref ✓
[INFRASTRUCTURE]FORGEML // LOG

Model Registry Over Hardcoded Paths

The FastAPI serving layer isolates model lifecycles behind an in-memory registry singleton (get_model). This decouples request routing from tensor loading, guarantees zero per-request initialization overhead, and allows seamless fallback between local weights and HuggingFace Hub.

Standard VerifiedSection Ref ✓
LIVE REST INFERENCE

Test The Model Inline

Calls POST /predict on the live FastAPI backend. No simulated latency or cached results.

SUMMARY TELEMETRY25,000 TEST OBSERVATIONS
91.30%
AccuracyDistilBERT test set (25k reviews)
+2.10pp
vs BaselineStatistically verified delta
15.0ms
InferenceSingle prediction on Apple Silicon MPS
5.90×
Batch SpeedupBatch-of-32 parallel forward pass