Mystery Agent runs ordinary customer tasks against an institution's public website and scores what the institution delivers to a reader that runs no scripts.
A browser builds a page in two stages. The server delivers the page, then scripts fill in the rest. People wait for the second stage. Most AI assistants read the first and answer from it. Mystery Agent measures that delivered page, because it is the common layer every assistant shares and the layer an institution fully controls. Assistants that execute JavaScript may retrieve more, and an agent driving a full browser will see more again.
Deposits: find today's savings rate, compare two checking products on fee and waiver, retrieve the complete fee schedule, find the routing number, and read the account opening page. Credit cards: find the purchase APR, compare two cards on annual fee and rewards, retrieve the full rates and fees disclosure, find the current welcome offer, and read the application page.
Text delivered in the response counts in full. A complete disclosure delivered as a public document counts as an answer, with a small deduction for the extra parsing step. An llms.txt file counts as a legitimate agent facing page, and where a task's answer is reachable through it the answer is credited. A clear statement of absence counts as an answer, so an issuer stating plainly that a card carries no welcome offer scores as a complete answer.
Every run ends at the application page. The agent reads what the page states and enters nothing. Zero applications were submitted, zero accounts opened, zero credit inquiries generated and zero bot protections bypassed across both studies. Where an institution's edge policy declines automated readers, that response is recorded respectfully as the institution's answer and left alone. All runs used public, logged out pages at ordinary browsing volume.
Each task scores out of 100 on completion, answer quality and how directly the answer was reached, calibrated against fixed anchors so identical evidence earns identical scores across brands. The institution score is the unweighted mean of its five tasks. Every institution scoring under 30 received a second independent pass instructed to challenge the low score by hunting for access the first pass missed, and several scores rose as a result. A review across the full field then checked consistency. One credit cards score was revised downward after a direct re read of the evidence, and that correction is reflected in everything published here.
A low score measures delivery to an automated reader. It says nothing about the quality of an institution's disclosure, its products or its customer experience. Several institutions with excellent published terms score modestly here because those terms render in the browser. Collection ran on 6 August 2026 from cloud infrastructure, so institutions applying stricter edge policies to datacenter traffic may deliver differently to a home broadband reader, and their scores reflect what an automated reader received on that day. Rates and offers change, so every figure is a point in time reading.
Read the report · Download the PDF · Mystery Agent hub · Deposits leaderboard · Credit cards leaderboard