A company we work with discovered last month that their Contact Us page had been replaced. Not updated by a developer, not accidentally overwritten, replaced. Anyone who clicked on it landed on a gambling site.
Their marketing team had no idea. Their agency (us) had no idea. We manage their LinkedIn presence, not their website, so nobody from the content side had clicked that page in some time. The website looked normal from the homepage. Navigation worked. The blog was live. The about page loaded cleanly. Only the Contact Us link was wrong, and it had probably been wrong for a while.
Nobody knows exactly how long.
Website compromises have always been an IT and security concern. What's changed is that they're now a content problem too, specifically an AI-search problem.
When ChatGPT, Perplexity, or any other LLM-backed tool crawls your website, it doesn't know which pages you approved and which were injected. It indexes what's there. If an injected page is on your domain and publicly accessible, any citation the model forms is a citation of that content, attributed to you.
A gambling redirect is visible enough that a human would notice immediately. The more dangerous version is subtler: a doorway page that shows your normal content to a logged-in user but serves a spam-optimised version to a crawler. Your team never sees it. The model does. The citation it returns, the next time someone asks a category question in your industry, points to your domain, and the content behind it isn't yours.
No GEO tool on the market currently checks for this.
The compromise we encountered was a single page, Contact Us, with a redirect injected at some point after the site was last properly audited. The attacker didn't touch anything visible enough to trigger an alert. The homepage ranked normally. The LinkedIn content we were posting continued to link cleanly to the main domain. Nothing looked wrong until someone clicked a specific link.
The lesson isn't that WordPress is insecure, or that small companies get targeted. It's that content teams, the people responsible for what a brand publishes, have no standard practice for checking whether the website they're pointing audiences toward actually contains what they think it contains.
A content audit means reviewing what you've written. A content-integrity audit means checking whether what's on your domain is actually what you wrote.
Those are different things, and right now almost nobody is running the second one.
These take under an hour across a typical site.
1. Click every page in your main navigation yourself. Not via a dashboard. Not via a broken-link tool. Actually click each URL and watch where it loads. A redirect won't show up as a 404, it resolves, just to the wrong place. The gambling site our client's Contact Us page had been pointing to would have passed any standard uptime check. The only way to catch it was to actually go there.
2. Run a site: search on Google. Type site:yourdomain.com into Google search and look at the results. You're not checking rankings, you're checking the count and the page titles. If Google is indexing 847 pages and you've published 200, something has been added that you didn't put there. If you see page titles in a language you don't publish in, or topics entirely unrelated to your business, those pages exist on your domain regardless of whether you can find them through your CMS.
3. Compare your sitemap to your CMS. Your XML sitemap lists every URL the site is asking search engines to index. Export it and cross-reference it against the pages actually in your CMS. Any URL in the sitemap that doesn't correspond to a page you created is worth investigating immediately. Attackers often add doorway pages that are crawlable but not linked anywhere visible, and the sitemap is how you find them.
4. Check your outbound links. Injected content frequently includes outbound links to the sites it's trying to boost. A free crawling tool will show you every external link your site is sending visitors to. If you see domains you don't recognise, especially clusters of links pointing to the same unfamiliar site, that's a flag.
5. Check what the crawler sees, not just what you see. Cloaking works by serving your normal page to a browser and a different page to a crawler, so view-source in your own browser won't catch it. Use Google Search Console's URL Inspection tool to view the live tested page as Googlebot renders it, or fetch the URL with a tool that identifies itself as Googlebot in its user agent, and compare that output line by line against what you see when you visit the page normally. Look for hidden text, hidden links, or content blocks that only appear in the crawler's version. This is the check that catches what the other four can miss.
The honest answer is that most content teams will say this is the developer's job. Most developers will say the site is the client's responsibility. Most clients will say they thought someone was watching it.
The gap is real, and it's not going away as more content gets published and fewer teams have dedicated site managers.
What changed in 2026 is the stakes. A compromised page was previously a brand safety and SEO problem. It's now also an AI-citation problem: if an LLM has crawled your injected page, the incorrect content is already in its context for your domain. Cleaning the page doesn't necessarily clean the citation.
Running these five checks takes less than an hour. We'd suggest making it a quarterly item on whatever content calendar you're already maintaining, alongside the checks you run on meta descriptions, broken images, and outdated posts.
Your content strategy depends on the assumption that your domain contains your content. It's worth verifying that that's still true.
The incident described in this piece involved a client whose website LexiConn does not manage. Details have been anonymised.
A content-integrity audit checks whether the pages on your domain are actually the pages you published. A conventional content audit reviews the quality of what you wrote; an integrity audit tests whether that is still what a visitor or a crawler receives. The two answer different questions, and most teams only run the first.
Because a language model crawling your site cannot tell an approved page from an injected one. It indexes what is publicly accessible on your domain, so a citation formed from injected content is still attributed to you. Removing the page later does not guarantee the citation goes with it.
A doorway page serves your normal content to a logged-in human and a spam-optimised version to a crawler. Because your own browser sees the correct version, view-source and visual checks pass. Only a crawler-perspective fetch reveals the difference.
Quarterly is a reasonable default, added to the content calendar alongside the checks you already run on meta descriptions, broken images and outdated posts. The full pass takes under an hour on a typical site.
In practice, nobody, which is the problem. Content teams treat it as a developer responsibility, developers treat the site as the client's, and clients assume someone is watching. Naming an owner is the first fix; the checks themselves need no technical skill.
Need expert content support? LexiConn has been India's B2B content partner since 2009, building content systems for leading enterprise brands across BFSI, technology, and media. Explore our content operations audit →