Model Evaluation
The systematic measurement of a model's quality, robustness, safety, latency, resource use, and generalization across representative data and operating conditions.
Machine-Learning Context
Model evaluation should cover not only average task quality but also robustness, safety, latency, resource use, calibration, and generalization across representative operating conditions. Offline benchmarks, curated golden sets, adversarial tests, human evaluation, and production telemetry provide different kinds of evidence.
Measurement Boundary
A single benchmark score does not represent production behavior. Data leakage, test contamination, distribution shift, and failure slices can make a high aggregate score hide important weaknesses.
Related Machine-Learning Concepts
- Word Error Rate
- Hallucination
- Latency
- Benchmark