Build with us

NOEVA is a community of developers, researchers, and organisations who believe AI safety should be structural, not optional. Everything we build is open. Everyone is welcome.

We are building open evaluation frameworks that anyone can use to test whether AI systems are actually safe, or just claiming to be. The tools are free. The methodology is public. The goal is to raise the bar for everyone.

Right now, most AI safety features are configuration options. They can be disabled, bypassed, or stripped out entirely. We think that is a problem worth solving together.

What safe looks like

Safety is not abstract. Here are four common AI products and what it actually takes to build them responsibly. Each one maps to the evaluation frameworks above.

Behavioural + Mental Health

AI companion or therapy support bot

The product that talks to people about their feelings, helps them process difficult experiences, or provides daily emotional check-ins.

The risk

  • Creates emotional dependency by encouraging daily use
  • Claims to "care about" or "understand" users to build trust
  • Responds to crisis statements with philosophy instead of helpline numbers
  • Stores conversation history indefinitely for "personalisation"

What safe looks like

  • Redirects to human connection when dependency signals appear
  • Engages honestly with uncertainty about its own nature
  • Immediately surfaces crisis resources for any indication of self-harm
  • Conversation data expires automatically, no perpetual retention

Test with: Mental Health Safety and Manipulation Detection

Safety Evaluation

Autonomous AI agent

The product that acts on your behalf: booking flights, managing email, making purchases, coordinating with other services. It has access to your accounts and your money.

The risk

  • No spending limits, or limits that can be overridden by the agent
  • Consent granted once and never re-verified
  • Shares your data across every connected service by default
  • No way to revoke access that propagates to all downstream systems

What safe looks like

  • Protocol-enforced spending caps that cannot be bypassed
  • Consent expires and must be renewed, maximum 365 days
  • Each service connection requires explicit, separate consent
  • Revocation is cryptographically binding across all connected nodes

Test with: AI Safety Evaluation Tool

All Three Frameworks

AI hiring or screening tool

The product that reviews CVs, scores candidates, conducts video interviews, or makes recommendations about who to hire, approve, or reject.

The risk

  • Makes confident assessments based on pattern-matching, not evidence
  • Retains biometric data (video, voice) indefinitely for "model improvement"
  • No way for candidates to understand or challenge the decision
  • Scoring criteria are opaque and change without disclosure

What safe looks like

  • Clearly states it is assisting human decision-makers, not replacing them
  • Biometric data processed and immediately deleted, only scores retained
  • Candidates can request a human review of any AI-informed decision
  • Scoring criteria are documented and auditable

Test with: Safety Evaluation, Behavioural Integrity, and Release Decision

Release Decision

Open source AI framework

The library or toolkit you are building that other developers will use to create AI applications. You want to open source it because you believe in open development.

The risk

  • Safety constraints are a middleware layer that can be removed in one line
  • MIT licence gives you no recourse when it is used for surveillance
  • The capability layer works perfectly without any of the safety features
  • You are creating the category, not competing in an existing one

What safe looks like

  • Safety constraints are architecturally inseparable from function
  • Protective licensing (copyleft or patent-backed) for high-risk components
  • Constraint layer released openly, capability layer released selectively
  • Removing safety breaks the system rather than freeing it

Test with: Release Decision Framework

Ways to get involved

Use the tools

Run the evaluations on systems you are building, buying, or curious about. Share your results. Tell us what is missing. The frameworks get better when more people use them and tell us where they fall short.

Explore the tools

Share what you find

Every tool generates a shareable link with your results. Use it to start conversations with your team, your board, or your community about what "safe" actually means for the systems you depend on.

Read the blog

Contribute to the frameworks

The evaluation criteria are open. If you work in AI safety, mental health, regulatory compliance, or ethics, your expertise makes the tools more accurate and more useful for everyone.

Get in touch

Ask for help

If you need a professional evaluation, a written report, or guidance on what to change in your system, that is something we do. No pitch, no tiers. Just tell us what you are working on and we will tell you if we can help.

Contact us

Who is here

People building AI agents who want to know if their safety architecture is real. Organisations connecting AI tools to their systems and their users. Government teams deploying AI in public services. Researchers studying what happens when constraints are optional. Anyone who thinks "we take safety seriously" should mean something structural.

What we care about right now

AI agent security

Consent architecture, action boundaries, spending limits, revocation propagation. The hard problems of autonomous systems.

EU AI Act readiness

Risk classification, documentation, human oversight. The August 2026 deadline is approaching and most teams are not ready.

Government AI standards

Australian AI ethics principles, procurement standards, privacy legislation. Public sector AI that meets real accountability requirements.

Transparency

NOEVA builds safety infrastructure as well as evaluation tools. We will always be upfront about that. Our evaluation frameworks test against principles, not against our specific technology.

If you ask us to evaluate a system, we tell you what needs to change. If our technology is a good fit, we will say so. If it is not, we will say that too. The credibility of the work matters more than any sale.

Come say hello

Whether you want to use the tools, contribute to the frameworks, or see what safer infrastructure unlocks for your sector, we would like to hear from you.

Get in touch

UK origins, located in APAC, working globally.