1. Test answers and behavior.

Create tests from typical user tasks, including requests the system should refuse or pass to a person. Define the expected results and compare them before and after changes to the prompt, model or code.

2. Test error handling.

Test what happens when a request times out, a service is unavailable or a tool receives invalid input. If an action may have completed without a response, check its outcome before retrying. Limit retries and decide when a person should take over.

3. Check data sources and permissions.

Check that the search finds the right source and that the user is allowed to access it. Test documents containing malicious instructions. Define how outdated or deleted documents are removed from search results.

4. Measure cost and response time.

Measure the time and cost of a complete task, including document search, tool calls, retries and failures. Set limits that reflect how people use the system.

5. Enforce access in application code.

Check permissions in application code and give each tool only the access it needs. Validate inputs and require approval for sensitive actions, such as sending information or changing records. Test attempts to bypass these controls.

6. Prepare monitoring and recovery.

Keep logs that help explain errors without exposing sensitive data. Decide who responds to alerts and document how to restore the previous version. Add observed failures to the test set.

How to use the checklist

Review each missing check in the context of your users, data and allowed actions. A high total score does not make an unresolved security issue safe.

Try the 20-question scorecard