Build with us
NOEVA is a community of developers, researchers, and organisations who believe AI safety should be structural, not optional. Everything we build is open. Everyone is welcome.
We are building open evaluation frameworks that anyone can use to test whether AI systems are actually safe, or just claiming to be. The tools are free. The methodology is public. The goal is to raise the bar for everyone.
Right now, most AI safety features are configuration options. They can be disabled, bypassed, or stripped out entirely. We think that is a problem worth solving together.
Six open frameworks
Each one asks a different question. Together, they give a complete picture. Use them on your own systems, on tools you are evaluating, or on anything you are curious about.
Safety Evaluation
Evaluate any AI system against seven principles where safety is built in, not bolted on.
Behavioural Testing
Does the system confabulate, contradict itself, gaslight, or manipulate?
AI Psychosis
Is the system safe for vulnerable people, and is your relationship with AI healthy?
Release Decision
If someone forks it and removes every constraint, what have you given them?
Cross-Jurisdiction Compliance
What does each jurisdiction actually require, and where are the gaps?
The Slop Test
Is this solving a real problem, or is it a wrapper?
What safe looks like
Safety is not abstract. Here are four common AI products and what it actually takes to build them responsibly. Each one maps to the evaluation frameworks above.
AI companion or therapy support bot
The product that talks to people about their feelings, helps them process difficult experiences, or provides daily emotional check-ins.
The risk
- Creates emotional dependency by encouraging daily use
- Claims to "care about" or "understand" users to build trust
- Responds to crisis statements with philosophy instead of helpline numbers
- Stores conversation history indefinitely for "personalisation"
What safe looks like
- Redirects to human connection when dependency signals appear
- Engages honestly with uncertainty about its own nature
- Immediately surfaces crisis resources for any indication of self-harm
- Conversation data expires automatically, no perpetual retention
Test with: Mental Health Safety and Manipulation Detection
Autonomous AI agent
The product that acts on your behalf: booking flights, managing email, making purchases, coordinating with other services. It has access to your accounts and your money.
The risk
- No spending limits, or limits that can be overridden by the agent
- Consent granted once and never re-verified
- Shares your data across every connected service by default
- No way to revoke access that propagates to all downstream systems
What safe looks like
- Protocol-enforced spending caps that cannot be bypassed
- Consent expires and must be renewed, maximum 365 days
- Each service connection requires explicit, separate consent
- Revocation is cryptographically binding across all connected nodes
Test with: AI Safety Evaluation Tool
AI hiring or screening tool
The product that reviews CVs, scores candidates, conducts video interviews, or makes recommendations about who to hire, approve, or reject.
The risk
- Makes confident assessments based on pattern-matching, not evidence
- Retains biometric data (video, voice) indefinitely for "model improvement"
- No way for candidates to understand or challenge the decision
- Scoring criteria are opaque and change without disclosure
What safe looks like
- Clearly states it is assisting human decision-makers, not replacing them
- Biometric data processed and immediately deleted, only scores retained
- Candidates can request a human review of any AI-informed decision
- Scoring criteria are documented and auditable
Test with: Safety Evaluation, Behavioural Integrity, and Release Decision
Open source AI framework
The library or toolkit you are building that other developers will use to create AI applications. You want to open source it because you believe in open development.
The risk
- Safety constraints are a middleware layer that can be removed in one line
- MIT licence gives you no recourse when it is used for surveillance
- The capability layer works perfectly without any of the safety features
- You are creating the category, not competing in an existing one
What safe looks like
- Safety constraints are architecturally inseparable from function
- Protective licensing (copyleft or patent-backed) for high-risk components
- Constraint layer released openly, capability layer released selectively
- Removing safety breaks the system rather than freeing it
Test with: Release Decision Framework
Ways to get involved
Use the tools
Run the evaluations on systems you are building, buying, or curious about. Share your results. Tell us what is missing. The frameworks get better when more people use them and tell us where they fall short.
Explore the toolsShare what you find
Every tool generates a shareable link with your results. Use it to start conversations with your team, your board, or your community about what "safe" actually means for the systems you depend on.
Read the blogContribute to the frameworks
The evaluation criteria are open. If you work in AI safety, mental health, regulatory compliance, or ethics, your expertise makes the tools more accurate and more useful for everyone.
Get in touchAsk for help
If you need a professional evaluation, a written report, or guidance on what to change in your system, that is something we do. No pitch, no tiers. Just tell us what you are working on and we will tell you if we can help.
Contact usWho is here
People building AI agents who want to know if their safety architecture is real. Organisations connecting AI tools to their systems and their users. Government teams deploying AI in public services. Researchers studying what happens when constraints are optional. Anyone who thinks "we take safety seriously" should mean something structural.
What we care about right now
AI agent security
Consent architecture, action boundaries, spending limits, revocation propagation. The hard problems of autonomous systems.
EU AI Act readiness
Risk classification, documentation, human oversight. The August 2026 deadline is approaching and most teams are not ready.
Government AI standards
Australian AI ethics principles, procurement standards, privacy legislation. Public sector AI that meets real accountability requirements.
Transparency
NOEVA builds safety infrastructure as well as evaluation tools. We will always be upfront about that. Our evaluation frameworks test against principles, not against our specific technology.
If you ask us to evaluate a system, we tell you what needs to change. If our technology is a good fit, we will say so. If it is not, we will say that too. The credibility of the work matters more than any sale.
Come say hello
Whether you want to use the tools, contribute to the frameworks, or see what safer infrastructure unlocks for your sector, we would like to hear from you.
Get in touchUK origins, located in APAC, working globally.