Certification independence
Delve Was the Loud Version. The Structural Version Is Still Legal.
The DeepDelver report identified something the compliance industry has been avoiding: when the same entity implements and examines, the compromise is structural, not incidental.
Analysis
The most consequential sentence written about AI compliance this year did not come from a regulator, an attorney, or a policymaker. It came from an anonymous former Delve customer, published on Substack under the pseudonym DeepDelver:
“By generating auditor conclusions, test procedures, and final reports before any independent review occurs, Delve places itself in the role of both implementer and examiner. This is not a technicality. It is a structural fraud.”
Delve, the Y Combinator-backed AI compliance automation startup that raised $32M at a $300M valuation, later parted ways with Y Combinator after the DeepDelver investigation surfaced. The alleged specifics were dramatic: fabricated evidence of board meetings that never happened, pre-generated auditor conclusions written before any review occurred, 99.8% identical reports across hundreds of clients, shell “US-based auditors” allegedly fronts for offshore certification mills, intellectual property allegedly taken from a fellow YC company. Delve has denied the allegations and stated that it does not issue compliance reports directly. Hundreds of Delve’s clients — including firms processing HIPAA-protected health data — nonetheless now hold documentation whose evidentiary weight, in a future regulatory or civil proceeding, is at minimum uncertain.
The scandal was covered as a story about one rogue startup that moved too fast and cut too many corners. That framing misses what the DeepDelver report actually revealed.
What Delve is accused of doing loudly, the compliance industry often does quietly.
The structural problem DeepDelver identified is not unique to Delve. It is endemic to the way AI compliance is currently sold.
Consider what happens when a Big 4 consulting firm sells an enterprise both AI governance advisory services and the audit that certifies whether the governance is adequate. The same firm helped the client build the compliance program. The same firm then attests to whether the compliance program is adequate. The financial relationship between the two engagements creates structural dependence on the ongoing client relationship. The auditor and the implementer are not the same person, but they work for the same P&L. That is a softer version of the same problem DeepDelver named.
Consider what happens when an AI safety consortium writes a certification standard, includes major AI companies among its authors, and profits from an insurance product backed by the certification. The entities being certified are the same entities whose input shaped the certification criteria. Downstream financial exposure through the insurance product creates additional dependence on the certification passing. The credit rating agencies had the same structure in 2007. The AAA ratings on subprime mortgage-backed securities were possible because the agencies rating the securities were paid by the entities issuing them, with downstream financial exposure through complex derivative structures. That structure looked authoritative until it was tested. Then it collapsed.
Consider what happens when an AI certification vendor sells a monitoring platform to the entities they audit, sells regulatory consulting to those same entities, and sells executive AI advisory alongside their assessments. Every subscription renewal, every consulting engagement, every advisory relationship depends on the certification not creating friction. The vendor’s independence from AI model vendors becomes irrelevant when the vendor has become financially dependent on the certified organizations.
Delve is accused of automating the compromise. Other models reproduce the structural risk more quietly. The 99.8% identical reports were Delve’s tell. Remove that specific tell, and the incentive structure across the compliance industry is the same one Delve is alleged to have exploited.
The regulatory precedent is already written.
Sarbanes-Oxley Section 201 addressed exactly this structural problem in the financial audit context after Enron. Congress prohibited registered public accounting firms from providing specific non-audit services to their audit clients; bookkeeping, financial information systems design and implementation, appraisal services, actuarial services, internal audit outsourcing, management functions, and legal services among others. The prohibition was not about auditor character. It was about auditor incentive. Congress recognized that when the same firm sold consulting to the client and then audited the results of that consulting, the audit’s evidentiary weight in future litigation or regulatory action was compromised regardless of the auditor’s intent.
That precedent is going to migrate into AI certification within the enforcement cycles now beginning. Illinois SB 315 mandates independent third-party audits, with the audit requirement taking effect January 1, 2028, or 90 days after a developer first qualifies as a large frontier developer, whichever is later. The statute also explicitly bars retention of any auditor where either party has a financial interest in the other, structural independence written directly into law. California SB 53, the New York RAISE Act, and Colorado’s AI Act are enacted, with operational obligations already in effect or approaching. The EU AI Act’s high-risk deadlines are staggered, December 2, 2027 for stand-alone systems, August 2, 2028 for product-embedded systems. Every one of these frameworks is likely to produce enforcement activity within twelve to eighteen months of its effective date.
The first enforcement action that tests what “independent third-party audit” legally means is going to examine the certifier’s financial relationship with the certified organization. The certifier that sold consulting to the client before the audit, the certifier that sold a monitoring product embedded in the client’s operations, the certifier that profits from an insurance product backed by their own certification standard, all of them will discover that the certification they issued is not the defensible evidence their clients thought they were purchasing.
Delve’s clients are the first to discover this. They will not be the last.
What actually holds up.
The certification model that survives regulatory and tribunal scrutiny has three structural properties, and all three must hold simultaneously.
First, no downstream financial interest in the certification outcome. The certifier does not sell consulting to the certified organization. The certifier does not sell monitoring products to the certified organization. The certifier does not sell insurance backed by their own certification. The certifier does not sell advisory services, remediation, readiness support, or any other product that creates financial dependence on the ongoing client relationship. Certification is the entire product, priced accordingly.
Second, human professional judgment at every audit decision point. No component of the audit, evidence gathering, framework mapping, control testing, finding determination, or certification decision, is delegated to an automated system or an AI agent. Automation may support delivery infrastructure. Automation does not perform the audit. When a regulator later asks who made the judgment that the evidence was sufficient, the answer is a qualified human auditor exercising professional judgment, not a model that cannot defend its reasoning in a courtroom.
Third, contemporaneous records that document the audit as it occurred, not as it was reconstructed. The evidentiary weight of contemporaneous records is orders of magnitude higher than the weight of narratives assembled after enforcement begins. Certifications that will survive litigation are certifications where the audit was documented in real time by a party with no financial interest in the finding.
An AI certification meeting all three criteria will be defensible when tested. A certification missing any one of them may not be.
The two-year question.
The Delve scandal was the visible failure of a business model that the compliance industry has been operating quietly for years. What made Delve exceptional was the crudeness of the execution, not the underlying structure. The structure is legal, widespread, and about to be tested by enforcement actions in every major AI regulatory framework in force.
Enterprises deploying AI in regulated industries have roughly two years to move their certification posture from convenience to defensibility. The organizations that align with certifiers meeting all three structural criteria before enforcement begins will have audit evidence that holds up when regulators or plaintiff’s counsel start asking questions. The organizations that align with certifiers who sell them consulting alongside the audit will hold documentation of good intentions when enforcement begins.
Delve’s clients did not choose fraud. They chose speed and convenience over defensibility, and the vendor they trusted delivered fraud dressed as convenience. The clients now facing potential HIPAA and GDPR exposure are collateral casualties of a certification model that was structurally compromised from the beginning.
The next Delve will be quieter. The certifications it issues will look defensible until they are tested. The clients who purchased them will discover the same thing DeepDelver’s employer discovered, on the same timeline the regulatory frameworks are creating.
There is only one way to build a certification model that survives that discovery. The industry knows what it is. The question is who will build it before enforcement forces the answer, and who will keep selling comfortable certification until it collapses under the weight of the first serious test.
Delve was the loud version. The structural version is still legal. That will not last.