Enterprise Software Due Diligence: Where AI-Built EdTech and SaaS Apps Fail the Review

Enterprise Software Due Diligence for AI-Built EdTech
Reported Incident

Security researcher Taimur Khan identified critical flaws in a Lovable-hosted EdTech application featured on the platform's Discover showcase, according to a February 2026 report by The Register. As reported by The Register, Khan demonstrated that an unauthenticated attacker could access 18,697 user profiles, modify student test scores, and wipe user accounts.

A security review is designed to identify risks such as broken authorization, tenant-isolation failures, exposed credentials, and inadequate operational controls. If you built your product fast with AI tools, that review is where you're most exposed.

Enterprise software due diligence is the structured evaluation an enterprise buyer conducts before executing a contract. Designed for EdTech and SaaS founders, this guide outlines what school districts, higher education institutions, and corporate buyers evaluate, highlights common vulnerabilities in AI-generated software, and provides actionable steps to bridge security gaps.

Haven't tested your app yet? Start with our vibe-coded app security checklist, then come back here.

The Six Dimensions

What do buyers check during enterprise software due diligence?

Buyers check six dimensions: financial stability, legal and regulatory standing, security posture, technical architecture, operational resilience, and reputational risk. That's the model Technology Match lays out for IT leaders, and each dimension fails in a different way.

The guide also warns that a valid SOC 2 certificate can sit on top of an architecture that buckles in production.

The Six Dimensions, and Where AI-Built Apps Are Exposed
Dimension What the buyer asks for Our read: risk for an AI-built app
Financial stability Audited financials, long-term viability, customer concentration (flagged above 40% of revenue from one customer) Low. It's a business question.
Legal and regulatory Signed DPA, FERPA and COPPA alignment, litigation and sanctions checks Medium. Practices must match the DPA.
Security posture SOC 2 Type II scope and exceptions, independent pen test from the last 12 months, incident response plan High
Technical architecture Data flows, SAML 2.0 or OIDC SSO, SCIM provisioning, audit logs exportable to a SIEM, tested RTO and RPO Highest. All three incidents below were access-control failures.
Operational resilience Uptime math, support tiers, notice before breaking changes Medium
Reputational risk Breach history, fourth-party dependencies, customer references Medium

Education items and risk column are our adaptation.

Why your product lands in Tier 1

Technology Match's framework puts any vendor that handles PII, PHI, or cardholder data, or that integrates deeply with core systems, into Tier 1. Tier 1 gets the full process: architecture review, a proof-of-concept in the buyer's own environment, and security exhibits in the contract.

Student records contain PII. If your product touches rosters, grades, or student work, assume you're Tier 1 from the first sales call.

What the education paperwork commits you to

Higher ed institutions use the HECVAT, the vendor assessment EDUCAUSE maintains with Internet2 and REN-ISAC, and HECVAT 4 launched in February 2025 with new privacy and AI questions. We cover what your SOC 2 report does and doesn't answer in HECVAT vs SOC 2 for EdTech vendors.

K-12 districts that use the Student Data Privacy Consortium's National Data Privacy Agreement (NDPA) ask vendors to sign one standard agreement instead of a custom contract. Version 2 was developed by 28 state alliances, and several of its clauses are engineering requirements in legal language.

NDPA v2.1 Clauses as Engineering Requirements
NDPA v2.1 clause What it requires of the provider What your system must be able to do
1.1 and 4.2
Purpose and use
Act as a School Official under the district's direct control, using data only for the contracted services Keep student data out of anything outside the contracted service
2.3
Subprocessors
Bind every subprocessor to terms no less stringent than the DPA Know every service that touches student data, including AI APIs
4.6
Disposition
Dispose of student data within 60 days of a request or termination Find and delete one district's data completely
5.2
Security audits
Run a security audit or assessment at least yearly and after any breach Produce a current report on request
5.3
Data security
Adopt a recognized framework such as NIST CSF, ISO 27000, or CIS Controls Map your controls to that framework
5.4
Data breach
Notify the district within 72 hours of confirming a breach Detect and scope an incident fast

Source: NDPA v2.1 standard text. Check the current version before signing.

The Research

Why do AI-built apps pass the demo but fail the review?

AI-built apps fail review because working code and secure code are separate properties, and today's models reliably deliver only the first. Veracode's 2026 GenAI Code Security Report found models write syntactically correct code nearly 100% of the time. Their average security pass rate is 56%, barely changed from 55% in Veracode's first report.

~100%
Of the time, models write syntactically correct code
56%
Average security pass rate, up from 55%
68%
Best 2026 model (GPT-5.5) — still fails one in three
Definition

Vibe coding is building software by describing what you want to an AI tool and shipping what it produces, with little or no line-by-line review. AI-assisted engineering uses the same tools, but a developer reviews, tests, and owns every change.

Buyers can't see which one you did. They can only see the evidence you hand them.

The average hides where the risk sits:

Security Pass Rate by Weakness Class
Weakness (CWE) Average security pass rate, 2026
Cryptographic algorithms (CWE-327) 87%
SQL injection (CWE-89) 83%
Cross-site scripting (CWE-80) 15%
Log injection (CWE-117) 12%

Veracode found models handle repeatable patterns but fail where they must trace how user input moves through an application. Model choice won't close the gap either. The 2026 leader, GPT-5.5, scored 68%, so it still fails nearly one in three security tasks.

A demo exercises the path a user takes. A security review exercises everything around it: a second tenant's data, a request edited in transit, a user deprovisioned last week.

Failure Record

What happens when nobody checks beneath the UI?

Data leaks through the same controls a buyer's reviewer asks about. All three incidents below are access-control failures, and each maps to a question you'll face in a real review.

Three Incidents, Three Review Questions
Incident What was exposed Root cause Review question that catches it
EdTech app on Lovable's showcase (Feb 2026) 18,697 user records, including 4,538 student accounts Authentication logic that blocked logged-in users and let anonymous visitors in "Show us unauthenticated requests failing on every endpoint."
Moltbook (Feb 2026) 1.5 million API tokens, 35,000 email addresses, private messages Database key in client-side JavaScript with no Row Level Security policies "How is tenant isolation enforced at the data layer?"
CVE-2025-48757 (May 2025) PII, API keys, and payment and subscription records in Lovable-generated projects Missing or insufficient Row Level Security, rated CVSS 8.26 "Prove one customer can't read another customer's rows."

The EdTech case hits closest to home. Users came from institutions including UC Berkeley and UC Davis, and Khan told The Register that K-12 schools with minors were likely on the platform. The Next Web described the most severe bug as inverted authentication logic that gave anonymous users full access.

Moltbook shows why "we use Supabase" isn't an answer. Wiz explained that the public key is safe only with Row Level Security policies in place. Without them, it gave anyone read and write access to the production database.

Definition

Row Level Security (RLS) is a database feature that filters which rows each user can read or change, based on who's asking.

Matt Palmer's CVE-2025-48757 disclosure notes that Lovable-generated apps call the database straight from the browser and rely exclusively on RLS, so one missing policy leaves nothing else in the way.

More failure patterns: why AI-generated apps break in production.

Framework

What is the Claim-vs-Proof Matrix?

Definition

The Claim-vs-Proof Matrix maps each reviewer question to the claim that won't pass and the evidence that will. Reviewers don't grade claims. They grade artifacts they can verify.

Use it as a self-audit. Count the rows where you could attach the proof to an email today. Every other row is a finding.

The Claim-vs-Proof Matrix
Buyer's question Claim that won't pass Proof that does
Single sign-on "Users sign in with an email and password." SAML 2.0 or OIDC integration tested against the buyer's identity provider
User provisioning "Admins can remove users." SCIM endpoint that creates and deactivates accounts from the identity provider, with test logs
Tenant isolation "Each school only sees its own data." RLS or equivalent enforced at the data layer, plus cross-tenant test results
Authorization "Protected pages require a login." Server-side checks on every endpoint, with unauthenticated and cross-role test results
Audit logging "We log everything." Documented log format and retention period, exportable to the buyer's SIEM
Penetration testing "We ran a scanner." Independent third-party pen test from the last 12 months, with an executive summary
Data deletion "We delete data on request." A documented, tested deletion procedure that meets the DPA deadline
Breach response "We'd let you know right away." A written breach response plan that supports 72-hour notification
AI and subprocessors "We use a major AI provider." Subprocessor list, signed terms, and a statement on whether customer data trains shared models
Recovery "We have backups." RTO and RPO documented and tested under failure conditions

Proof column draws on Technology Match's Tier 1 checklist and NDPA v2.1.

Two rows matter most in EdTech. Deletion carries a hard deadline, since the NDPA sets a 60-day window. AI is new territory, and HECVAT 4 added a dedicated AI question set for it.

Want a second set of eyes on your matrix?

Request a readiness review and we'll tell you which rows a reviewer will flag before a buyer does.

Request a Readiness Review →
Implementation

How do you get an AI-built app ready for enterprise due diligence?

You get ready by producing the Proof column. Each step ends in an artifact you can hand a buyer.

  1. Map where the data lives. List every place student or customer data goes, including logs, analytics, file storage, and AI provider calls. The artifact is a data flow diagram and data schedule.
  2. Move authorization to the server and the database. On Postgres or Supabase, enable RLS on every table and scope each policy to the user and tenant. Palmer notes RLS works on full rows, so sensitive columns belong in separate tables or behind server-side access.
  3. Test like the reviewer. Automate unauthenticated, cross-tenant, and deprovisioned-user requests, and run them on every deploy. The artifact is a passing test report.
  4. Build enterprise identity. Add SAML 2.0 or OIDC single sign-on, SCIM provisioning, and role-based access control. The artifact is integration documentation for the buyer's identity team.
  5. Make audit logs exportable. Record who accessed what and when, set a retention period, and support SIEM export. The artifact is a log schema and retention policy.
  6. Get independent verification. Commission a third-party pen test, fix the findings, and plan SOC 2 Type II if enterprise sales are on your roadmap. Our SOC 2 guide for EdTech startups covers timing and scope.
  7. Write the documents reviewers ask for. Prepare a breach response plan, deletion procedure, subprocessor list, and AI data-use statement.

Steps 2 and 3 matter most for AI-built code, because they target the controls behind all three incidents above.

Contract Stage

How do due diligence findings turn into contract terms?

A gap you can't close before signing doesn't always kill the deal. Technology Match advises buyers to map every residual risk to a contract clause, so your gap becomes your obligation.

That can mean breach-notification deadlines, audit rights, limits on training or secondary use, and liability super-caps for data breaches.

Every row you evidence before signing is a clause you don't carry afterward.

FAQ

Frequently asked questions

Enterprise software due diligence is a buyer's structured review of a vendor before signing. It covers six dimensions: financial stability, legal and regulatory standing, security posture, technical architecture, operational resilience, and reputational risk. Vendors that handle PII, including student records, get the deepest review, with architecture checks and a proof-of-concept.

Not on its own. A SOC 2 Type II report is strong evidence for security posture, but buyers still review technical architecture, identity integration, logging, and data handling separately. In education, institutions can also require a HECVAT or a signed data privacy agreement, and each asks questions SOC 2 doesn't answer.

Tests only check what they're written to check. If the same AI session writes the feature and its tests, the tests can inherit the feature's assumptions and only confirm the normal user path. A reviewer tries what those tests skip: an anonymous request, another tenant's data, or a user who was already removed.

For a Tier 1 review, expect the question. Technology Match's checklist asks which SSO protocols a vendor supports, such as SAML 2.0 or OIDC, and how SCIM handles user provisioning. SSO lets the buyer control sign-in through its own identity provider, and SCIM creates and removes accounts automatically as people join or leave.

Technology Match's checklist flags a SOC 2 Type I report offered in place of Type II, expired or narrowly scoped evidence, and no independent pen test in the last 12 months. It also treats vague answers as a signal in themselves. For K-12, add a deletion process that can't meet the NDPA's 60-day window.

Yes. HECVAT 4, released in February 2025, added a dedicated set of AI questions, and IT due diligence checklists now ask whether customer data trains shared models. Expect to document which AI providers process data, what they receive, whether they retain it, and where those limits appear in your contracts.

Strategic Summary

The bottom line

Enterprise software due diligence rewards proof, not polish. An AI-built product can pass the demo and still fail the review, because the review tests what the demo never touches: tenant isolation, identity, logging, and deletion.

The fix isn't to stop building with AI. It's to produce the Claim-vs-Proof evidence before a buyer asks.

US-Led. Cebu-Powered.

Ready to find out what a reviewer would flag in your product?

Hireplicity is a US-led engineering team based in Cebu that builds and secures EdTech and SaaS platforms. Book a 30-minute strategy call and we'll walk through your matrix, row by row.

Book a 30-Minute Strategy Call →
References
  1. Technology Match, "IT Vendor Due Diligence: A Practical Process and Checklist for IT Leaders," published September 9, 2025 — https://technologymatch.com/blog/how-to-do-vendor-due-diligence-as-an-it-leader
  2. Veracode, "2026 GenAI Code Security Report: AI Is Writing More of Your Code but Security Hasn't Caught Up," July 28, 2026 — https://www.veracode.com/blog/2026-genai-code-security-report-ai-risk/
  3. Access 4 Learning / Student Data Privacy Consortium, "National Data Privacy Agreement (NDPA) Standard Version 2.1" — https://files.a4l.org/privacy/NDPA/NDPA_v2-1_STANDARD_WEB.pdf
  4. Student Data Privacy Consortium, "National Data Privacy Agreement" — https://privacy.a4l.org/national-dpa/
  5. EDUCAUSE Review, "HECVAT 4: Better than Ever," February 10, 2025 — https://er.educause.edu/articles/2025/2/hecvat-4-better-than-ever
  6. The Register, "AI-built app on Lovable exposed 18K users, researcher claims," February 27, 2026 — https://www.theregister.com/software/2026/02/27/ai-built-app-on-lovable-exposed-18k-users-researcher-claims/5038511
  7. The Next Web, "Lovable security crisis: 48 days of exposed projects, closed bug reports, & the structural failure of vibe coding security," April 21, 2026 — https://thenextweb.com/news/lovable-vibe-coding-security-crisis-exposed
  8. Wiz, "Hacking Moltbook: AI Social Network Reveals 1.5M API Keys," February 2, 2026 — https://www.wiz.io/blog/exposed-moltbook-database-reveals-millions-of-api-keys
  9. Matt Palmer, "CVE-2025-48757," May 29, 2025 — https://mattpalmer.io/posts/2025/05/CVE-2025-48757/
Next
Next

How to Vet a Software Development Partner With a Paid Trial Sprint