AI Accessibility: What Designers Flag on AI-Built Landing Pages
Sidharth Nayyar

AI can now produce a landing page that looks finished at first glance. What it still gets wrong is mostly about reading. Over 2026, Contra Labs had panels of working designers mark up pages built by frontier models. Layout, spacing and hierarchy was the top flag across the studies, and on Claude Opus 5 pages typography and color contrast came next: the same things that decide whether a page is usable for people with low vision, people who zoom, and people on small screens. That makes AI accessibility less about the model and more about the review step after it. We set their findings against our own scan of 1,079 websites, where color contrast is the most common specific WCAG failure.
Key findings
- On Claude Opus 5 pages, the top three issue tags were layout, typography, and color contrast.
- 63% of the rated issues on those pages needed rework before a client could see them.
- Every page in the July study drew marks. The model with the fewest marks still averaged 3.4 per page, per designer.
- In ad images, typography grew from 3% to 34% of designer comments in the final round.
- In our own scan of 1,079 websites, color contrast failed on 69.3% of sites, the most common specific WCAG failure we found.
What designers flag on AI-built landing pages
In its August study, Contra Labs asked nine designers to review 20 landing pages: five client briefs, each built by four models. Pages were shown one at a time, in random order, with no model name. Designers boxed each problem, tagged it by type, and rated it Minor, Major, or Blocker.
Claude Opus 5 drew the most marks, 208, and its top three tags all concern how readable and scannable the page is.
Issue tags on Claude Opus 5 landing pages
Number of designer annotations carrying each tag. One annotation can carry several tags.
| Category | Value |
|---|---|
| Layout, spacing and hierarchy | 88 |
| Typography | 69 |
| Colour and contrast | 58 |
| Brand fit and tone | 11 |
| Originality or generic feel | 8 |
Source: Contra Labs, “What stands between Opus 5 and client-ready pages: layout and readability”, August 7, 2026. Base: 208 annotations, 5 pages, 9 designers.
Most were not cosmetic notes. Of the 183 Opus annotations with a severity rating, 86 were Major and 29 were Blockers. For the other three models the rework share ran from 48% to 53%, so no model was clean.
The tags also stack on the same element. A faint button label is a type problem and a contrast problem at once. Typography and color contrast appeared together on 31 annotations, and 23 of those were rated Major or Blocker.
63% of rated issues on Opus 5 pages needed rework before showing a client (115 of 183). Contra Labs, August 7, 2026. Percentage is our calculation.
74% of issues tagged both typography and contrast were Major or Blocker (23 of 31). Contra Labs, August 7, 2026. Percentage is our calculation.
Two pages show where this bites. Ashfall, a dark landing page for an indie game, won praise for its palette and big display type. Designers still flagged supporting copy and button labels that sank into the background. On Switchboard, text in the how-it-works timeline overlapped and the pricing cards lost their alignment. Those are the sections a visitor reads to understand the product and compare plans, and a review that stops at the hero would miss them.
Are AI website builders accessible?
These studies test model output, not builder products, so they cannot rank an AI website builder for accessibility. What they do show is that the first output is a draft. In the July study, eight designers reviewed every page from four models across five briefs, and every page drew marks.
Problems marked per page, per designer
Average annotations each designer left on each landing page, by model.
| Model | Annotations per page |
|---|---|
| Muse Spark 1.1 (210 annotations) | 5.25 |
| Claude Fable 5 (191 annotations) | 4.7 |
| Grok 4.5 (184 annotations) | 4.6 |
| GPT 5.6 Sol (136 annotations) | 3.4 |
Source: Contra Labs, “Where four AI models break when they build a landing page”, July 16, 2026. Base: 5 briefs, 8 designers, 40 designer-page reviews per model. GPT 5.6 Sol was reviewed in a separate session.
Volume is only half the picture. GPT 5.6 Sol drew the fewest marks, but 37% of its tags were layout, spacing, and hierarchy, and typography added another 15%. Claude Fable 5 had the lowest share of layout tags and more polish and consistency marks per page than any other model. Muse Spark 1.1 drew the most marks, and 48% of them were Major or Blocker. A lower count does not mean a page is accessible. It means a different list of things to fix.
One more distinction matters for buyers. These studies looked at the pages a model produced. None of them tested whether a person using a screen reader can operate the builder itself.
Why the last round of AI web design is about type and usability
The April Human Creativity Benchmark followed work through three rounds: ideation, mockup, and refinement. Different models won different rounds. In Contra Labs’ later website evaluations, GPT 5.6 Sol led ideation, Claude Opus 5 led mockup, and Claude Fable 5 led refinement. No model led every stage.
What changed between rounds was where designers looked. In ad images, typography went from about 3% of comments early on to about 34% in refinement. In landing pages it more than doubled, from about 3% to about 7%, and usability became the top theme. Designers also agreed with each other most in that final ad round, on legible type, clear calls to action, and contrast.
Usability worked like a gate. Low-scoring images rarely made the top two, however good they looked.
How often an ad image finished in the top two, by usability score
Usability scored by designers on a 1 to 5 scale. The study does not publish the figure for score 4.
| Usability score | Finished in top two |
|---|---|
| Score 1 | 10% |
| Score 2 | 22% |
| Score 3 | 36% |
| Score 4 | Not published |
| Score 5 | 84% |
Source: Contra Labs, “The Human Creativity Benchmark”, April 30, 2026. Base: Human Creativity Benchmark ad-image evaluations, 5 evaluators per domain.
The practical lesson: check text and controls after the visual direction is settled, not before. A page that already looks finished is exactly when muted supporting copy slips through.
How AI and accessibility connect: designer flags mapped to WCAG
Contra Labs asked designers what looks and works right, not whether pages meet an accessibility standard. None of the studies measured contrast ratios or tested with a screen reader, so these are design judgments, not WCAG failures.
But the categories designers flag most are the ones WCAG 2.2 already measures. A designer spots a faint label by eye. A checker reports the same label as text below 4.5:1 contrast. The table below is our mapping.
| Designer tag | WCAG criteria | What to check |
|---|---|---|
| Colour and contrast | 1.4.3 Contrast (Minimum); 1.4.11 Non-text Contrast | Body text at 4.5:1 or more, large text at 3:1. Button borders, input outlines and icons at 3:1 against what sits next to them. |
| Typography | 1.4.4 Resize Text; 1.4.12 Text Spacing | Faint, thin supporting copy is the usual failure. Zoom to 200% and apply the text-spacing overrides: nothing may clip or overlap. |
| Layout, spacing and hierarchy | 1.3.1 Info and Relationships; 1.4.10 Reflow | Visual hierarchy must also be real hierarchy: one h1, heading levels in order, lists as lists. At 320px wide, no sideways scrolling. |
| Interaction and affordances | 2.4.7 Focus Visible; 2.5.8 Target Size (Minimum) | Tab through the page: every control shows a focus ring. Tap targets at least 24 by 24 CSS pixels, or spaced so a 24px circle around each does not overlap another. |
Not every layout note is an accessibility issue. Some are taste. But the readability notes usually are, and they affect more people than designers: anyone on a phone in sunlight, older readers, anyone who zooms. Where the ADA or the European Accessibility Act applies, a page that fails 1.4.3 is also a legal risk. For a deeper look at the two most common flags, see our guides to color contrast accessibility and accessible font size.
What our scanner finds on real websites
Designers judged these pages by eye. We can add a measured view. In August 2026 we published our State of Web Accessibility 2026 report: 1,079 websites scanned against WCAG 2.1 and 2.2 Level A and AA. The criteria that test each of the four designer flags fail on half of those sites or more.
Sites failing the WCAG check behind each designer flag
Share of 1,079 websites with at least one failure of the criterion. Designer flag in brackets.
| WCAG criterion (designer flag) | Sites failing |
|---|---|
| 1.3.1 Info and Relationships (layout) | 78.6% |
| 1.4.3 Contrast (color) | 69.3% |
| 2.4.7 Focus Visible (interaction) | 67.1% |
| 1.4.12 Text Spacing (typography) | 50.5% |
Source: WebAbility, “State of Web Accessibility 2026”, August 19, 2026. Base: 1,079 websites, scanned 21 July to 18 August 2026.
Color contrast is the clearest overlap. It failed on 69.3% of sites and made up 21% of every issue we detected, the largest single share. That is the same problem designers flagged by eye on AI-built pages: supporting text and button labels that fade into the background.
The two data sets measure different things. Contra Labs recorded designer judgments on AI-built pages. Our report counts automated WCAG failures on sites that ran through our scanner, built with all kinds of tools. Neither proves the other. Both point to the same place: text that people cannot read. Automated engines file many structural issues under 1.3.1, so that bar is a broad category, not one defect. That is why our report names contrast, at 69.3%, as the most common specific failure.
How to check an AI-generated landing page before launch
- Scan the page with our free website accessibility checker to catch contrast, heading, and label failures.
- Check the text and background pairs the model picked with a color contrast checker. Look hardest at muted supporting text, button labels, and text on images.
- Zoom to 200% and narrow the window to 320px. Nothing should clip, overlap, or scroll sideways.
- Tab through the page. Every link and button needs a visible focus ring and a clear name.
- Check the heading outline. The hierarchy the model drew must match the real h1 to h3 order.
Review the whole page, including pricing and the sections below the fold, and check again after the last design edit. Our WCAG compliance checklist covers the wider review. Automated tools catch roughly 30 to 40 percent of WCAG issues, so they are a first pass, not the whole check. A manual accessibility audit covers what scanners miss.
AI accessibility FAQ
Are AI website builders accessible out of the box?
Treat the first output as a draft. In the Contra Labs annotation studies, designers marked problems on every AI-built landing page they reviewed. Layout and hierarchy led, and typography and color contrast were common. Scan the page, fix what the scan finds, and test with a keyboard and zoom before launch.
What is the most common accessibility problem in AI web design?
Readability. Layout and hierarchy was the top designer flag across the studies. On Claude Opus 5 pages, typography and color contrast came next. Faint supporting text and button labels that blend into the background came up repeatedly. Those map to WCAG 1.4.3 contrast, which an automated checker can measure. In our own scan of 1,079 websites, contrast was the most common specific failure, on 69.3% of sites.
Can AI-generated websites meet WCAG?
Yes, but rarely on the first output. Designers flagged problems on every AI-built page in the Contra Labs annotation studies. Treat generated code as a draft: scan it, fix contrast, headings, labels and focus, then test with a keyboard, zoom and a screen reader before launch.
Did the Contra Labs studies test WCAG compliance?
No. Designers judged how pages looked and worked. The studies did not measure contrast ratios, test with screen readers, or check WCAG conformance. The WCAG mapping in this article is our analysis of which success criteria test the problems designers flagged.
Which AI accessibility tools should I run on a generated page first?
Start with an automated scan for contrast, headings, form labels, and link names, since those match the issues designers flag most. Then check zoom to 200%, a 320px-wide layout, and keyboard focus by hand. A manual audit covers screen reader and context issues that tools miss.
About the data
All study figures come from Contra Labs and are linked below. Percentages marked as our calculation are derived from their published counts. The samples are small: five to nine designers per study and one output per model per brief, so a single strong or weak page can move the totals. Agreement between designers was limited across the website evaluations (Krippendorff’s alpha of 0.21). Read these as strong signals about where AI design breaks, not as model rankings.
The website figures come from our own State of Web Accessibility 2026 report. It covers 1,079 sites scanned with our current engine from 21 July to 18 August 2026. These sites ran through our scanner, so they are not a random sample of the web. The report publishes its full method and limits.
- The Human Creativity Benchmark, Contra Labs, April 30, 2026
- Where four AI models break when they build a landing page, Contra Labs, July 16, 2026
- What stands between Opus 5 and client-ready pages: layout and readability, Contra Labs, August 7, 2026
- How model performance shifted across four website design evals, Contra Labs, August 26, 2026
Quick Questions
Tap to ask AI about this article







