There was a time when evaluating an AI system meant running a test set, computing an accuracy score, and calling it done. That was good enough when models answered questions in isolation. It is not good enough anymore.
Need help?
Contact usThere was a time when evaluating an AI system meant running a test set, computing an accuracy score, and calling it done. That was good enough when models answered questions in isolation. It is not good enough anymore.