Evals are how you measure the quality and effectiveness of your AI system. This is a new term that I see trending a bit. An eval tells if you got the solution right or not. A simple as that on the surface.
I find this progression very logical. LLMs and AI cannot be strictly controlled to exact precision. There will always be an element of surprise in the output and outcomes.
I've for long wanted to hear admissions of not knowing and making errors in product work. To me, this has always been human. Alas, many of the real-life work settings have still been about knowing and being right - always. We've kind of pretended to know more than we really know. This has been so in engineering. The same requirement of knowing it right has been in product management too. With AI, that expectation of a priori certainty is no longer there. Now we can really embrace the iteration, failing initially and then correcting ourselves.
(Side note: the certainty was never there even without AI and LLMs. We just often pretended we know and get everything right by ourselves.)
Evals come in many forms: Metrics and telemetry; LLM-based second checks and evaluations; and of course Human in the loop watching. The common denominator is investing into sensing for correctness of the product in real use.
Genie and the lion tamer
In my imagination, I've had this metaphor of genie and the lion tamer as the duality of AI. The AI is the genie. Just like in the Aladdin movie the genie is a bit wonky and out of control. The genie has an abundance of ideas - too many. Many of the ideas don't really fit as good solutions. You certainly can't trust the genie.
The lion tamer is the control of the genie. He cracks his whip and keeps the genie in control. He watches and brings the show back to order.
Evals is the lion tamer. Actually not quite. Evals is the eyes of the lion tamer. Evals notices when things are not going right. Then we need somebody to do something about the situation.
Investing in being in control - Having a system of being eventually in control
Make the effort to know what is right and wrong in the business. AI can generate anything - both right and utterly wrong. It is still your responsibility to tell what is right and wrong. In a way, to tell what is real and what is hallucinations.
Make the investment proactively. Put in the sensing with and for the customers. Then you can have AI search for an optimal solution. Evals is the sensing.
You cannot know everything a priori. That foresight just does not exits. That foresight certainly doesn't exist the complex word where we are living in. The future is unpredictable and often surprising. Predictions will be wrong. This however does not mean that we could not understand and sense. The uncertainty does not mean that anything goes. There are righter and wronger solutions. It is for you to know the difference.
Not knowing beforehand is OK. Not putting the effort to understand and gain knowledge is not OK. I welcome the shift to more exploratory approaches. To me, this is just one more step to more agility.
Now what are evals in practice? That's for you to imagine and build. Looking forward to hearing about your novel approaches.
How do you approach evaluating AI systems in your work? What metrics matter most to you?