Quality Assurance Labs
AI Apps & Integration

Computer Vision in Production — The Testing Checklist

Senior AI Engineer7 min readPublished Updated

A CV model that hits 95% accuracy in the lab can drop to 60% in production. Here's the checklist we use to catch that before users do — lighting, edge devices, bias, latency, drift, and adversarial inputs.

Camera inspecting objects with detection frames
#computer-vision#edge-AI#model-testing#ML-QA#TensorFlow-Lite

Computer vision models are notoriously bad at generalizing. A model trained on clean data fails on messy inputs. A model that hits 95% accuracy in your test set can collapse to 60% in the field.

Here's the testing checklist we use at QA Labs before any CV model ships.

Test against real-world conditions

Lab accuracy ≠ field accuracy. Test against:

  • Real-world lighting (not studio)
  • Different camera angles
  • Motion blur
  • Occlusion (partial obstructions)
  • Weather conditions (for outdoor)
  • Different times of day

Test on edge devices

Latency and accuracy change dramatically on edge hardware:

  • Measure inference time on target devices
  • Check memory usage under load
  • Verify battery impact
  • Test model quantization (if applicable)

A model that runs fine on GPU servers might time out on a Raspberry Pi.

Test for bias

CV models often perform worse on underrepresented groups:

  • Measure accuracy across demographics
  • Check for false positives/negatives by group
  • Test with diverse datasets
  • Audit training data for representation

Bias testing is non-negotiable for any CV model used in decisions affecting people.

Test for model drift

Model accuracy degrades over time:

  • New patterns appear
  • Data distributions shift
  • Environments change

Plan for periodic retraining. Set up monitoring for accuracy drops.

Test adversarial inputs

CV models can be fooled by:

  • Adversarial perturbations (small pixel changes)
  • Out-of-distribution inputs
  • Trick images

For security-sensitive applications, adversarial testing is required.

What we typically find

  • Models trained on studio data fail on real-world images
  • Latency on edge devices exceeds SLA
  • Demographic bias in face-related models
  • Accuracy drops 10–20% over 6 months without retraining

Key takeaways

  • Test CV in real conditions, not just the lab
  • Edge devices expose latency and accuracy issues
  • Bias testing is non-negotiable
  • Plan for drift and retraining
  • Adversarial testing for security-sensitive deployments

Further reading

About the author

Senior AI Engineer →

Senior AI Engineer · Quality Assurance Labs

Notes from the lab.

Testing, engineering and growth — delivered to your inbox.

Need CV model QA? Book a scoping call

Let's talk →