Enforcing Accessibility Standards Consistently Across Mobile Builds
Automated checks miss real barriers that only assistive-tech users encounter.

Accessibility standards fail in mobile development because teams treat enforcement as an event instead of a constraint. An audit happens, a review happens, a pre-release check happens, and in between those moments, nothing holds the standard in place. Problems accumulate quietly during every sprint that falls between one inspection and the next, and surface all at once in the next audit, when someone finally looks.
The pattern repeats across product cycles with little variation. A pre-release audit turns up a list of violations. The team fixes them under deadline pressure, often in the final days before a submission. The next build cycle starts with no mechanism watching for the same mistakes, and within a few sprints, the same classes of violations are back: missing accessible names, broken focus order, contrast ratios that don't hold up under VoiceOver.
None of this traces back to teams not caring or not knowing the rules. Most engineers can describe what WCAG requires, what VoiceOver compatibility means in practice, and what touch target sizing demands on a small screen. What's missing is a process architecture that turns that knowledge into something the pipeline actually checks on every build. The standard exists in a document somewhere. It does not exist as a constraint the system enforces. That gap, between what people know and what the pipeline checks, is where accessibility breaks down, and it's the gap every section that follows is built to close.
How AI coding agents make periodic enforcement structurally untenable
AI coding agents write code faster than any human reviewer or periodic audit process can verify it, and accessibility suffers more than most other quality dimensions because unit tests rarely catch it. A broken button handler throws an error. A missing accessible name just sits there, silently unusable to a screen reader, until someone using one happens to hit it.
Research out of UC Berkeley describes how these agents actually work: an agent iterates on code until some evaluator accepts the result as done. The evaluator checks requirements against a model of the deployment environment, not against the environment itself, and both the requirement and the model are only approximations. Neither one captures the full intent of the people who asked for the work, and neither one captures what will actually happen once real users touch the product. A requirement like "this screen must be navigable by a VoiceOver user" sits precisely in the space between what an automated evaluator is built to check and what a real user actually experiences. An evaluator can confirm an element has a label. It cannot confirm that label makes sense when read aloud in sequence with everything around it.
Handing an agent the WCAG documentation does not close this gap. Access to a standard and adherence to a standard are two different things, and enforcement is supposed to fill the distance between them. Carnegie Mellon and Stanford researchers add a second complication: agent-generated patches run consistently longer than the patches a human would write to solve the same problem. More code means more surface area for a violation to hide in, and it means a reviewer has to do more work to reach the same level of confidence they'd have had with a shorter, human-written change. Put the two findings together: the old rhythm of writing code at human pace and auditing it on a schedule no longer holds. Code volume is rising. Verification capacity is not rising with it.
What automated accessibility checkers catch and miss
Detection tools do one job well: they find violations. They identify a missing label, a contrast failure, an ARIA role that was left unset. What they don't do is fix the problem, hold the fix in place across future builds, or tell anyone how a real assistive-technology user will actually experience the screen.
Research from the Technical University of Munich and Nanyang Technological University lays out this gap directly. Tools like IBM's Accessibility Checker and axe-core have gotten better at finding violations over time, but neither one repairs what it finds. Developers are still left to read the WCAG rule, interpret it, and write the fix by hand, which makes remediation slow, inconsistent from one engineer to the next, and prone to error under deadline pressure. The same research shows violations cluster rather than appear alone. A single interface typically carries several structurally related problems that need coordinated edits across multiple files, and the tools on the market today handle each flagged violation in isolation, without any sense of how one fix might conflict with another. The repair that results is incomplete at best and contradictory at worst. The researchers point to the Ant Design homepage as an illustration: a professionally maintained, high-profile open-source UI framework that still shows missing accessible names, insufficient contrast, unlabeled ARIA roles, and content sitting outside landmark regions, all at once, on a single page.
The deeper problem sits below static analysis. Research from UC Berkeley and the University of Michigan tested computer use agents under assistive-technology conditions, keyboard-only navigation and screen magnification, and found their performance dropped substantially compared to how the same agents performed under default sighted-user conditions. The agents simply don't reflect the way blind and low-vision users actually move through an interface. A build can pass every automated scan on record and still put up real barriers for someone navigating with a screen reader or a keyboard alone.
What this adds up to is a basic property of what static analysis can and cannot see, marking the line where detection stops and genuine enforcement has to begin.
Enforcing accessibility at the build level
Build-level enforcement starts from a different premise than a scan or a checklist. The standard gets defined once, expressed as a requirement an evaluator can check against every single build, and treated as a gate. A violation blocks the merge. It does not get filed away as a finding in a report that someone may or may not read before the next release.
Two distinct problems have to be solved for that gate to work, and the UC Berkeley research names both of them clearly. The first is the requirement gap: the distance between what a stakeholder actually wants and what the evaluator is told to check. Accessibility enforcement breaks down when the requirement written into the evaluator is vaguer than the real standard, something like "has an accessible name" standing in for what's actually needed, which is an accessible name that matches the design specification and means something to a VoiceOver user hearing it read aloud. Closing that gap takes a human decision about what "accessible" means for a specific screen, a specific flow, a specific interaction, written down precisely enough that a machine can check it consistently. That's a governance act. No amount of tooling configuration substitutes for it.
The evaluation environment itself has to approximate how a real assistive-technology user interacts with the interface, which introduces a second gap. The A11y-CUA research found that computer use agent performance under magnification conditions dropped to 28.3%, a sharp divergence from how the same agents performed under default interaction. That number alone makes the point that the environment most evaluators model is not the environment accessibility actually depends on.
Practitioners have started calling the response to this harness engineering: building the surrounding infrastructure, custom tests, monitoring, architectural constraints, orchestration, that governs how an AI agent is allowed to generate and modify code in the first place, so that accessibility sits inside the structure as a constraint rather than arriving as a check after the fact. None of this is meant to take the release decision away from a person. The Carnegie Mellon and Stanford position paper argues the next real advances in coding agents will come from designing for task alignment, verifiability, steerability, and adaptability between human and agent. An enforcement system that strips human judgment out of the release decision has just moved that judgment somewhere harder to see.
Apple's store review model and the limits of top-down enforcement
Apple's App Store review process is the clearest large-scale example of what top-down accessibility governance looks like once it's actually running. Apple encourages VoiceOver compatibility and Dynamic Type support, and reviewers may cite them during the review process, but these standards are not written in as universal rejection criteria that apply to every app. Apple's Accessibility Nutrition Labels are self-reported by developers themselves, and the published evaluation criteria read as guidance for indicating support rather than as a hard pass/fail gate every submission must clear.
Even with those limits, the mechanism has real consequences attached to it. Developers are encouraged, and will eventually be required, to supply specific information about their app's accessibility features in its metadata, and the review process is what turns a stated standard into something with teeth: a rejection, a resubmission, a delay. A gate with a consequence attached produces different behavior from a report full of recommendations, and Apple's review shows that clearly.
What it also shows is where that kind of gate falls short. Review happens at the end of the release cycle, after development is finished. It catches problems at the most expensive possible point: a rewrite under time pressure, a resubmission clock running, revenue sitting on hold while the fix gets made. A team that treats App Store review as its accessibility enforcement strategy still relies on the event-based model. It has pushed the event to the last possible moment in the process, where a violation costs the most to fix and the least time remains to fix it well.
QA headcount reductions and the urgency of automated enforcement
The industry is restructuring QA staffing in a way that removes the human fallback periodic accessibility audits have always depended on. For two decades, the standard staffing model assumed a reasonably large QA team manually exercising builds before release. That model is giving way to smaller teams of more senior quality engineers, backed by tooling that's supposed to cover the verification work the lost headcount used to do.
The logic behind that shift only holds if the tooling is actually in place before the headcount reduction happens, not after. Accessibility verification is one of the hardest categories to hand off to general-purpose tooling, because it depends on modeling how a real assistive-technology user interacts with an interface, not just running a static scan across the code. A scanner can tell a team their contrast ratios pass. It cannot tell them whether a VoiceOver user can actually complete the checkout flow.
The measurement gap makes this worse, not better. The UC Berkeley two-gap framework, paired with the broader argument about how much evaluators approximate rather than confirm, means a leader cutting QA headcount may have no reliable way to know whether accessibility is still being covered by whatever verification infrastructure remains. That's a visibility problem layered on top of a staffing problem.
Accessibility is a rejection vector at the app store level today, and an emerging legal compliance requirement for native mobile apps on top of that, not a quality dimension a team can push off to a final human review once QA capacity tightens. A missed violation under constrained QA capacity isn't a minor bug report. It appears as a rejected build, a resubmission cycle, or a compliance exposure that costs far more to resolve after the fact than it would have cost to catch at the point the code was written.
What consistent enforcement requires
Consistent enforcement depends on three things holding together at the same time: a standard precise enough for a machine to evaluate automatically, an evaluator that models real assistive-technology interaction instead of just checking static code properties, and a gate that attaches real consequences to a finding instead of letting it sit in a report nobody acts on.
The first of those is a human responsibility. The requirement gap identified in the agentic software engineering research closes only when people write down, precisely, what "accessible" means for a particular screen, a particular flow, a particular interaction pattern, in language specific enough for an automated evaluator to apply the same way every time. No scanner invents that precision on its own.
The second depends on closing the model gap. The A11y-CUA research shows that the interaction environment for blind and low-vision users is measurably different from how a default sighted user interacts with the same screen. Any enforcement system that doesn't account for that difference will underreport violations consistently, build after build, no matter how often it runs. That's why accessibility enforcement has to work across three layers at once: the product intent layer, where design decides whether an interaction is meant to be accessible in the first place; the implementation layer, where code either delivers that intent or doesn't; and the runtime layer, where the result either holds up or fails under real assistive-technology conditions.
None of this argues for taking engineers and reviewers out of the loop. The agents can run the checks at the speed code gets written. The decision to approve a release, to accept a trade-off, to judge whether a fix actually serves the user it's meant for, stays with the people who understand what the standard was written to protect.


