Armilla Joins Global Consortium Building Industry Certification Standards for AI Systems

July 21, 2026
5 min read

By Philip Dawson

As the first insurer in the consortium, Armilla is contributing an AI system evaluation and underwriting perspective to the shared assessment infrastructure that safe AI adoption depends on.

Last month, Armilla AI joined the Appia Foundation as a founding member, alongside OpenAI, Microsoft, Google, Mastercard, Ericsson, Schneider Electric, Siemens, Mitsubishi Electric, and other global organisations working to advance trusted certification and conformity assessment for AI systems.

For Armilla, this is both a milestone and a natural continuation of our work on AI assurance: building the infrastructure that trustworthy AI adoption depends on, and enabling more rigorous underwriting of AI risk and liability.

Appia is an international collaboration developing the conformity specifications that organisations can use to demonstrate that their AI system(s) meet the risk and compliance obligations that apply to them. Hosted by the Linux Foundation's Joint Development Foundation (JDF), Appia is designed to sit between the standards and regulations that set expectations for AI systems and the conformity assessments that verify whether those expectations have been met.

Standards and frameworks, like ISO/IEC 42001, NIST’s AI Risk Management Framework, Singapore's AI Verify, and the EU AI Act, have each helped define what responsible AI systems and organisations are expected to do. What the market still needs is credible, reusable, and interoperable ways to measure AI outcomes and demonstrate conformity with commercial and regulatory expectations, especially at the system level.

A Consortium Across the AI Value Chain

One of the most important features of the Appia Foundation is the breadth of its founding membership. The consortium spans the entire AI value chain by bringing together organisations that deploy AI across industry, energy, networks, and payments; foundation model developers, platform and compute providers; and organisations focused on testing, assessment, certification, and insurance.

Breadth matters, because conformity assessments cannot be developed in isolation from the market they are meant to serve. Enterprises, model providers, assessors, regulators, and insurers all need evidence (as produced by conformity assessments) that is technically meaningful, commercially practical, and recognised across the value chain.

Armilla joins as the first member from the insurance industry. This insurance perspective matters because underwriting depends on evidence. The more credible, comparable, and independently governed that evidence is, the more useful it becomes for pricing AI risk with confidence.

From Standards to Evidence

Appia's work turns obligations into specifications that can be assessed in practice. Its approach is organised across two layers: a Requirements and Guidance layer, which addresses what is required; and an Assessment Enablement layer, which addresses how those requirements are evaluated.

Many AI governance frameworks define the outcomes that organisations should pursue or the controls they should have in place. Yet, fewer provide the assessable, modular, and reusable criteria needed to demonstrate conformity in a way that the whole value chain can recognise.

For AI assurance to scale, evidence produced in one part of the market should carry forward to another part of the market. For example, a deployer should be able to rely on evidence from a supplier, a customer should be able to understand what has been assessed, a regulator should be able to see how obligations have been evaluated, and an insurer should be able to tell the difference between an unsupported assertion and a system assessed against credible criteria.

This is the practical role of conformity assessment. It does not replace standards, regulation, or governance frameworks. It makes them easier to apply, evaluate, and rely upon.

Why Insurance Belongs in the Conversation

Insurance has long depended on independent standards, testing, and certification to make new categories of risk legible. Underwriters Laboratories, founded in 1894 with support from the fire insurance industry, established a practical signal that electrical products conformed to recognised safety expectations. FM Global built a comparable philosophy into property insurance through engineering, loss prevention, and certification.

AI is not electrical equipment or industrial property. It is more dynamic, contextual, and probabilistic. AI systems can change after deployment, behave differently across use cases, and sit beneath many downstream applications. But the underlying market logic holds whereby standards create expectations, independent assessment produces evidence, and evidence makes risk easier to understand and underwrite.

AI insurance markets remain constrained in part because the methods used to evaluate AI systems have not always been designed with underwriting in mind. Benchmarks can be useful, but they rarely capture how a system behaves in operational conditions. Governance policies matter, but on their own they do not establish how a model performs, fails, changes, or is monitored over time. Underwriting works by reconstructing loss scenarios: the specific ways a system can fail, the conditions under which that failure occurs, and the losses that follow. Pricing those scenarios means putting numbers to them: how likely each failure is, and how severe its consequences would be. 

This is where measurement of AI systems becomes essential. Probability and severity cannot be read off a benchmark score, inferred from a governance policy, or extrapolated from litigation history, which today is largely concerned with previous generations of less capable, less autonomous systems. They have to be derived from evidence about how the system actually behaves in use. What insurers need, then, is evidence that speaks to those scenarios, is operationally relevant, and is produced through a process the market can trust.

Neutrality and Openness Matter

The governance of this work matters because AI assessment criteria should not belong to any single company, institution, government, or sector. For AI certification to support trust across the market, their specifications must be developed through an independent, neutral and open process.

This is why Appia's home within the Linux Foundation's JDF is significant. The JDF gives competitors and stakeholders from across the AI ecosystem a structure to contribute safely to shared work, with governance that supports openness, public scrutiny, and a path to broader standardisation.

For underwriting, this is not procedural. Evidence is more credible when it is grounded in specifications developed through a neutral, transparent, and independently governed process. Proprietary frameworks may be useful in particular contexts, but they are less likely to provide the common foundation that a broader AI assurance market needs. Neutrality, openness, and independence are part of what allows portability of trust across the value chain, including among peers, partners, and competitors.

Armilla's Perspective

Armilla's participation in Appia builds on work we have pursued since our founding. Long before AI insurance became a defined category, we were focused on the role of independent assessment in making AI systems more trustworthy. Our team contributed to early certification efforts through the Certification Working Group convened by the University of Toronto's Schwartz Reisman Institute, worked alongside organisations such as the Responsible AI Institute, and supported global assurance initiatives including Singapore's AI Verify.

Across these efforts our view has stayed consistent: trustworthy AI requires credible, independent evidence. This view has led us to connect assessments directly to risk transfer as better assessments produce better risk data, better risk data supports more precise underwriting, and more precise underwriting can reward the organisations that invest seriously in performance, governance, safety, and monitoring.

Looking Ahead

It is likely that no single standard, certification, or assessment scheme will capture everything that matters about an AI system. Different sectors, jurisdictions, applications, and risk profiles call for different forms of assurance. The market does not need one universal test. It needs a common basis on which multiple assessments and certifications can be developed, compared, and relied upon with greater consistency. And for certification to be trusted across markets, it also needs to be consistent with established global norms for conformity assessment, including the ISO/CASCO model that already underpins recognised certification in other industries.

Appia is an important step in that direction. As AI systems are deployed more widely in enterprise and regulated settings, organisations will need practical ways to assess them rigorously, certify them credibly, and insure them where appropriate. The task ahead is to build open, interoperable, and independently governed ways of turning AI expectations into evidence that customers, regulators, auditors, and insurers can rely on.

Share this post

Ready to Insure Your AI?

Armilla’s AI insurance is your fail-safe against fast-evolving AI risks. We combine deep technological insight with robust insurance solutions so you can focus on innovation, without interruption.
Get in Touch