THE RESULT
What the data shows
The largest computed metric is 0.3467 for Other_Faults; the smallest is 0.02834 for Dirtiness. Metric: sample_share (Share of labeled faults).
The denominator contains faulty observations only. This is not a defect rate across manufactured plates.

| fault | records | sample share |
|---|---|---|
| Other_Faults | 673 | 0.3467 |
| Bumps | 402 | 0.2071 |
| K_Scatch | 391 | 0.2014 |
| Z_Scratch | 190 | 0.09789 |
| Pastry | 158 | 0.0814 |
| Stains | 72 | 0.03709 |
| Dirtiness | 55 | 0.02834 |
THE METHOD
From source to answer
Class counts and proportions across recorded faults.
The Python source exposes this study’s transformations. The complete project download includes shared preparation and evaluation routines.
THE NEXT DECISION
What follows from the finding
Use class balance to choose metrics and inspect rare-class coverage before model fitting.
Where the conclusion stops
Every source row is a recorded fault. The data cannot estimate a production defect rate or distinguish healthy plates. Batch and machine IDs are absent, so grouped exact-input holdouts do not prove factory transfer.
Related studies may reuse observations or holdouts. These are historical analyses; associations and backtests do not demonstrate commercial impact. Further model tuning needs new, untouched evaluation data.