Hamfy research notes / 06

Reading plant-disease AI results without losing the test context

What evidence should accompany a claim that a plant-disease model is ready for use in a growing facility?

01 / Purpose

Scope of this review

A comparative reading of two primary research papers, supported by image-acquisition documentation and an AI evaluation framework. The studies concern different tasks and conditions. This is an evidence-boundary analysis, not a benchmark of current models or a Hamfy product test.

02 / Approach

How we assessed the evidence

We checked the PlantVillage figures against the original paper’s abstract and external-image discussion, and the cassava study against its field-evaluation methods and results. We kept the metrics and test populations separate. We did not retrain models, relabel images or pool the studies’ scores.

03 / Findings

What the evidence supports

The test population travels with the score

The PlantVillage paper supports a distinction between performance within a controlled collection and performance on a small external image set. Our practical conclusion is to require a dataset description beside any headline result.

Field presentation belongs in the evaluation

The cassava study makes symptom severity and capture mode explicit. This informed our recommendation to include difficult, representative observations instead of testing only obvious examples. It does not quantify the expected performance of an indoor canopy camera.

Acquisition and review need their own records

The PlantCV guidance informs capture planning, while NIST provides the general evaluation context. Hamfy’s proposed workflow adds image-quality checks, independent testing and a named reviewer for uncertain observations.

04 / Boundaries

What this review does not establish

  • The studies were published in 2016 and 2019 and are not a survey of the latest model capabilities.
  • The PlantVillage external sets are small web-image samples; their scores do not estimate accuracy in every farm.
  • The cassava evaluation uses a different crop, detection task and metrics, so no pooled score or direct model ranking is justified.
05 / Application

Hamfy’s proposed next step

Ask a proposed supplier to describe the task and the test conditions before discussing an accuracy target. Plan a site-specific evaluation with independent examples and a record of missed events, false alerts and review effort. Decide what claim the evidence supports before extending the system’s role.

Evidence to collect for a real project

  • Representative images from the intended installed camera view.
  • Labels with a documented confirmation method and an uncertainty category.
  • An independent test population matching the proposed operating scope.
  • An operator review log tied to the exact model and threshold evaluated.
06 / References

Source notes

  1. Peer-reviewed experiment · 2016Using Deep Learning for Image-Based Plant Disease DetectionMohanty, Hughes & Salathé · Frontiers in Plant Science

    Source of the three reported evaluation results. DOI: 10.3389/fpls.2016.01419. The external-image tests are described in the Discussion.

  2. Peer-reviewed field evaluation · 2019A Mobile-Based Deep Learning Model for Cassava Disease DiagnosisRamcharan et al. · Frontiers in Plant Science

    Provides a separate field example with still/video and symptom-severity comparisons. Its metrics are not directly interchangeable with PlantVillage accuracy.

  3. Technical documentation · v4.4Analysis approachesPlantCV documentation

    Supports acquisition planning around the intended observation and analysis.

  4. Public framework · 2023AI Risk Management Framework: CoreNIST

    Provides the general basis for evaluation in context and continued monitoring.

Back to the article