Original Research · 2026

CPU-Constrained Deep Learning
for Tomato Disease Detection

Obidur Rahman, Lipon Chandra Das, Arnab Aich, Abu Saiman Md Taiham, Atif Ibna Latif

Under Review, Springer Book Proceedings

Agriculture loses 40% of its yield to disease. GPUs power modern AI, but they stay out of reach for farmers in developing regions who rely on basic laptops. We benchmark ResNet-50, ConvNeXt-Tiny, and FastViT-T8 on consumer CPU hardware, looking for models that balance accuracy with real-world deployability.

Earlier studies ran their models on NVIDIA Tesla or RTX GPUs. Most farmers in South Asia and Africa can't buy that hardware; they have ordinary laptops and phones. We looked for a model that stays accurate and fast on a plain CPU, so disease detection can reach the 180M+ ton global tomato market.

We evaluated three architectures on an AMD Ryzen 5 5600G (6C/12T, no GPU).

ResNet-50

25.6M params · Baseline

Traditional CNN. Stable but computationally heavy for CPU inference.

ConvNeXt-Tiny

29.0M params · Modern CNN

Transformer-inspired architecture with 7×7 kernels. Highest parameter count.

FastViT-T8

4.03M params · Hybrid

CNN-Transformer hybrid. 6× smaller. Optimized for edge inference.

Batch: 8 (Eff: 32)Optimizer: AdamWLR: 1e-4 → 5e-5Cosine DecayInput: 224×224
PLACEHOLDER: architecture_diagram.jpg

PlantVillage subset, 16,012 images across 10 disease classes. Class imbalance ratio of 8.6:1 (Yellow Leaf Curl: 3,209 vs Mosaic Virus: 373). Standard 70/15/15 train/val/test split.

PLACEHOLDER: dataset_samples.jpg

FastViT-T8 gives the best balance of speed and accuracy: 99.66% accuracy at 0.022s per image (45 FPS), 57% faster than ConvNeXt-Tiny while giving up only 0.22% accuracy. ConvNeXt-Tiny reaches 99.88% but takes 0.051s per image.

ModelAccuracyPrecisionRecallF1Latency
ConvNeXt-Tiny99.88%0.9990.9980.9980.051s
FastViT-T899.66%0.9970.9960.9960.022s
ResNet-5097.69%0.9780.9760.9760.055s
PLACEHOLDER: benchmark_chart.png
PLACEHOLDER: confusion_matrix.png

A few honest caveats before you trust these numbers:

  • Data leakage. Images were split at random, not per plant. The reported accuracy is likely an upper bound.
  • One run. Results come from a single training run. More runs are needed to confirm the 0.22% gap is real.
  • Lab photos. PlantVillage was shot against plain backgrounds. Field photos will do worse.
  • Memorisation. ConvNeXt hit 100% training accuracy against 99.88% validation, a sign it memorised part of the data.

If this work is useful in your research, please cite:

@inproceedings{rahman2026cpu,   title={CPU-Constrained Deep Learning for Tomato Disease Detection: Traditional, Modern, and Hybrid CNN Comparison},   author={Rahman, Obidur and Das, Lipon Chandra and Aich, Arnab and Taiham, Abu Saiman Md and Latif, Atif Ibna},   booktitle={Springer Book Proceedings},   year={2026},   note={Under Review}, }