Multi-Model Evaluation & Benchmark Center
Live side-by-side evaluation across ConvNeXt-Tiny, Swin-Transformer, Multi-Modal Fusion, INT8 Quantized, EfficientNet-B0, and Rule Classifiers on holdout test datasets.
Overall Model Performance Leaderboard (Holdout Test Split)
Evaluated on 241 Test Samples| Architecture | Test Accuracy | Macro F1 | Weighted F1 | CPU Latency | Model Size | Performance Highlights |
|---|---|---|---|---|---|---|
| ConvNeXt-Tiny (Modern Pure CNN) | 92.4% | 0.915 | 0.92 | 28.5 ms | 27.8 MB | Production Winner (Best Real-World Generalization) |
| Calibrated Weighted Consensus (Optimal F1-Soft Vote) | 89.0% | 0.889 | 0.89 | 28.0 ms | 218.0 MB | Multi-Model Ensemble |
| Swin-Transformer (Self-Attention) | 91.7% | 0.917 | 0.918 | 34.2 ms | 28.2 MB | Best Context |
| INT8 Quantized Dynamic Engine | 84.2% | 0.842 | 0.844 | 12.4 ms | 4.8 MB | Fastest CPU |
| EfficientNet-B0 (Baseline Pure Vision) | 84.2% | 0.841 | 0.843 | 32.1 ms | 18.9 MB | Baseline |
No Images Found
No images matched the selected split and class filter.