Key idea: An offline test set only checks what you thought of. Watch real traffic, add what fails to the test set, and fix it so both scores agree.
Offline score vs. live score
v1v1 + failures in test setv2v2 + failures in test set
- Run v1. The offline test set says 8/8. What does the live week say?
- Look at the live misses in the results. What do they have in common?
- Turn on Add this week's failures to the test set and run v1 again. Does the offline score finally show the problem?
- Switch to v2 and run. Do the offline and live scores agree now?
- Which day looked worst on the dashboard? How long would you have shipped broken routing without it?