Here are the seven questions listed on the slide:
- What is your measured error rate, by language and by name origin?
- Can you reproduce a decision from six months ago — inputs, version, output?
- What can it do without a human, and where is that written down?
- Private data, untrusted input and outbound comms at once? Which did you remove?
- How do you test that it disagrees with the user when the user is wrong?
- What happens when the model underneath you is deprecated?
- What does it refuse to do?
Here are the seven questions listed on the slide: