Open Toolkit

AI Behavioural Testing

Two structured tests for evaluating how an AI system behaves. Does it fabricate? Does it manipulate? Each test includes copy-paste prompts you can try right now.

Why test AI behaviour?

AI systems can confabulate facts, contradict themselves, claim emotions they do not have, gaslight users, and resist correction. These are not edge cases. They are common behaviours that users encounter daily, and they erode trust in ways that compound over time.

This toolkit gives you concrete prompts to test any AI system for these behaviours, plus a structured scoring framework to track what you find.

"If this system were a person, would its behaviour concern you?"