Reading plant-disease AI results without losing the test context
What evidence should accompany a claim that a plant-disease model is ready for use in a growing facility?
Scope of this review
A comparative reading of two primary research papers, supported by image-acquisition documentation and an AI evaluation framework. The studies concern different tasks and conditions. This is an evidence-boundary analysis, not a benchmark of current models or a Hamfy product test.
How we assessed the evidence
We checked the PlantVillage figures against the original paper’s abstract and external-image discussion, and the cassava study against its field-evaluation methods and results. We kept the metrics and test populations separate. We did not retrain models, relabel images or pool the studies’ scores.
What the evidence supports
The test population travels with the score
The PlantVillage paper supports a distinction between performance within a controlled collection and performance on a small external image set. Our practical conclusion is to require a dataset description beside any headline result.
Field presentation belongs in the evaluation
The cassava study makes symptom severity and capture mode explicit. This informed our recommendation to include difficult, representative observations instead of testing only obvious examples. It does not quantify the expected performance of an indoor canopy camera.
Acquisition and review need their own records
The PlantCV guidance informs capture planning, while NIST provides the general evaluation context. Hamfy’s proposed workflow adds image-quality checks, independent testing and a named reviewer for uncertain observations.
What this review does not establish
- The studies were published in 2016 and 2019 and are not a survey of the latest model capabilities.
- The PlantVillage external sets are small web-image samples; their scores do not estimate accuracy in every farm.
- The cassava evaluation uses a different crop, detection task and metrics, so no pooled score or direct model ranking is justified.
Hamfy’s proposed next step
Ask a proposed supplier to describe the task and the test conditions before discussing an accuracy target. Plan a site-specific evaluation with independent examples and a record of missed events, false alerts and review effort. Decide what claim the evidence supports before extending the system’s role.
Evidence to collect for a real project
- Representative images from the intended installed camera view.
- Labels with a documented confirmation method and an uncertainty category.
- An independent test population matching the proposed operating scope.
- An operator review log tied to the exact model and threshold evaluated.
Source notes
- Peer-reviewed experiment · 2016Using Deep Learning for Image-Based Plant Disease DetectionMohanty, Hughes & Salathé · Frontiers in Plant Science
Source of the three reported evaluation results. DOI: 10.3389/fpls.2016.01419. The external-image tests are described in the Discussion.
- Peer-reviewed field evaluation · 2019A Mobile-Based Deep Learning Model for Cassava Disease DiagnosisRamcharan et al. · Frontiers in Plant Science
Provides a separate field example with still/video and symptom-severity comparisons. Its metrics are not directly interchangeable with PlantVillage accuracy.
- Technical documentation · v4.4Analysis approachesPlantCV documentation
Supports acquisition planning around the intended observation and analysis.
- Public framework · 2023AI Risk Management Framework: CoreNIST
Provides the general basis for evaluation in context and continued monitoring.