In May 2024, an automated bot called yoshi-code-bot pushed 2,596 modules of internal Google documentation to a public GitHub repository. The files described 14,014 attributes from Google’s Content Warehouse API — a detailed catalog of the signals, metrics, and systems Google uses to evaluate and rank web pages. The documents were never meant to be public. They stayed exposed for approximately six weeks before being removed.
SparkToro co-founder Rand Fishkin and iPullRank CEO Mike King were the first to publish detailed analyses after receiving the documents from Erfan Azimi, who later identified himself as the source. Google initially refused to comment. When press coverage became unavoidable, spokesperson Davis Thompson confirmed the documents were authentic but cautioned against “making inaccurate assumptions about Search based on out-of-context, outdated, or incomplete information.”
That caveat is worth taking seriously. The documents don’t reveal the weights Google assigns to each signal, don’t tell us which signals are actively used versus deprecated, and don’t explain how these signals interact within the ranking algorithm. What they do reveal is the full scope of data Google collects and considers — and in several cases, this scope directly contradicts years of public statements by Google engineers and spokespeople.
This article covers every major finding from the leak, explains what each one means for SEO practitioners, and identifies the specific contradictions between Google’s public position and its internal systems.
How the Leak Happened
The documents originated from Google’s internal code base and were exposed when developers included Content Warehouse API documentation in a code review that got pushed to a public GitHub repository. The timeline: the files appeared on GitHub around March 13, 2024, remained publicly accessible until approximately May 7, 2024, and were analyzed and published by Fishkin on May 27, 2024.
The documents cover API specifications — they describe the data structures, modules, and attributes that Google’s systems use to process and rank content. They don’t contain the actual ranking algorithm (the formulas, weights, and decision trees), but they catalog what inputs the algorithm consumes. Think of it as seeing every ingredient in a kitchen without knowing the exact recipe.
Google’s response was carefully worded. They acknowledged the documents were genuine but emphasized they might be “outdated” or “incomplete.” When pressed on specific contradictions with past public statements, Google said they don’t confirm or deny “sensitive” information about Search to prevent manipulation by bad actors.
Separately, the U.S. Department of Justice’s antitrust trial against Google in 2023-2024 independently confirmed several findings from the leak — particularly NavBoost and the use of click data in rankings — through testimony from Google’s own VP of Search, Pandu Nayak. This cross-verification through a legal proceeding under oath significantly strengthens the credibility of the leaked documents.
SiteAuthority: The Metric Google Said Didn’t Exist
For years, Google representatives publicly denied that Google calculates or uses any kind of domain-level authority metric. John Mueller, Gary Illyes, and other Googlers repeatedly stated that Google evaluates pages individually rather than assigning site-wide authority scores.
The leaked documents tell a different story. Within the CompressedQualitySignals module, there’s an integer field called siteAuthority. It functions as a persistent, composite score calculated at the site or sub-domain level. According to Mike King’s analysis, this metric serves as a foundational input for preliminary ranking in Google’s Mustang system.
This doesn’t necessarily mean Google’s siteAuthority works identically to the “Domain Authority” metrics calculated by third-party tools like Moz or Ahrefs. The calculation methodology is unknown. But its existence is unambiguous — Google does maintain and use a site-level authority signal in its ranking systems, despite publicly denying it.
What contributes to siteAuthority isn’t fully documented, but the surrounding attributes suggest it incorporates anchor signals, on-site prominence, content quality signals, link diversity, and Chrome user behavior data.
What this means for SEO
Building site-level authority matters — not just page-level optimization. This validates the long-held SEO practitioner view that a domain’s overall strength affects individual page rankings. Strategies that improve the entire site (consistent content quality, strong internal linking, accumulating authoritative backlinks) contribute to a rising siteAuthority score that benefits every page on the domain.
NavBoost: Click Data as a Core Ranking Signal
NavBoost is mentioned 84 times in the leaked documents. It’s a system that tracks user engagement with search results and uses that data to adjust rankings. The DOJ antitrust trial independently confirmed its existence and importance — Pandu Nayak testified that NavBoost is “one of the most important” ranking signals Google uses.
This directly contradicts years of public statements. Gary Illyes, Google’s search relations lead, dismissed the possibility of using clicks for rankings in multiple public appearances. Matt Cutts, Google’s former head of web spam, said the same. The leaked documents leave no ambiguity.
NavBoost tracks several specific click metrics:
goodClicks — clicks that result in a satisfied user (long dwell time, no return to search results).
badClicks — clicks where the user quickly returns to search results, indicating dissatisfaction (“pogo-sticking”).
lastLongestClicks — the final result a user clicked on and viewed for the longest duration in a search session. This is a particularly strong satisfaction signal — it suggests the user found what they were looking for and stopped searching.
unsquashedClicks — clicks deemed genuine and valuable (as opposed to bot traffic or manipulated clicks).
NavBoost operates on a rolling 13-month window of aggregated click data. It’s also region-specific — click patterns from users in one geographic region may influence rankings differently than patterns from another region.
The leaked documents also reference a “trendSpam” demotion for CTR manipulation, indicating Google has detection systems for artificial click inflation. Attempts to manipulate click signals may produce temporary ranking improvements, but the documents suggest these gains are ephemeral — rankings tend to revert once artificial clicks stop, and persistent manipulation risks demotion.
What this means for SEO
User satisfaction is a direct ranking input. Pages that consistently satisfy searchers (long engagement, no pogo-sticking) accumulate positive NavBoost signals over time. This makes genuine user experience optimization — page speed, content quality, answering the query thoroughly — a ranking strategy, not just a nice-to-have. It also means that a page ranking well today but failing to satisfy users will gradually lose its position as negative click data accumulates.
The Sandbox Is Real: hostAge and New Site Treatment
Google employees have denied the existence of a “sandbox” — a mechanism that restricts new websites from ranking well for an initial period — since the mid-2000s when SEOs first coined the term. Mueller’s response when asked about it was characteristically dismissive.
The leaked documents contain an attribute called hostAge within the PerDocData module. The accompanying notes state it’s used to identify and handle “fresh spam during serving time.” This confirms that Google does apply different evaluation criteria to new websites based on domain age.
The exact mechanics aren’t fully described. The documents don’t specify how long the sandbox period lasts or what criteria allow a site to emerge from it. But the existence of the attribute is unambiguous — new sites are subject to different signal processing than established sites.
What this means for SEO
New websites should expect a period of reduced visibility regardless of content quality or optimization. Launch strategies need to account for this delay. Building initial authority through brand mentions, social presence, and links from established sites during the sandbox period can help accelerate the transition. The 2024 antitrust trial documents suggest that link profile age and establishment history are factors in how quickly a site exits sandbox treatment.
Chrome Data: Google Tracks More Than It Admitted
Matt Cutts stated publicly that Google does not use Chrome data for rankings. John Mueller reaffirmed this claim. The leaked documents show the opposite.
Multiple modules reference Chrome-sourced data. A field called chromeInTotal captures site-level views from Chrome browsers. Additional attributes track Chrome user behavior including time on page, bounce rate, scroll depth, pages per session, and interaction events.
These signals feed into site-quality assessment at the aggregate level — not just individual page rankings. Chrome data appears to be used as a validation layer, confirming (or contradicting) the quality signals derived from crawled content and link analysis.
The scale of Chrome’s market share makes this significant. Chrome holds approximately 65% of global browser market share. Google has access to behavioral data from a massive sample of web users — data that no other search engine or third-party tool can replicate.
What this means for SEO
If Chrome users consistently bounce from your site within seconds, that behavioral signal feeds into Google’s quality assessment of your domain. This reinforces the importance of genuine user experience — not just technical SEO metrics, but actual on-site engagement. Sites where users stay, scroll, interact, and return build positive Chrome-sourced signals over time.
Content Freshness: Three Dates, One Assessment
The leak reveals that Google doesn’t rely on a single timestamp to assess content freshness. Three distinct date signals are tracked:
bylineDate — the date explicitly stated in the article’s byline or publication date field. This is what most content management systems output as the “published” or “updated” date.
syntacticDate — a date extracted from the URL structure or document title. If your URL contains “/2024/06/” or your title includes “2026 Guide,” Google captures that.
semanticDate — an estimated date derived from the content itself, its anchor context, and related documents. Google’s systems can infer when content was actually written based on the topics it discusses, the sources it references, and the temporal context of surrounding content.
This three-date system explains why simply changing a publication date without updating the actual content doesn’t fool Google. The semanticDate may still indicate the content is old if its references, data points, and examples haven’t been updated.
Content freshness feeds into the Query Deserves Freshness (QDF) algorithm, which boosts recent content for queries where timeliness matters — breaking news, recurring events, trending topics, and subjects that change rapidly. Not all queries trigger QDF. A search for “how photosynthesis works” doesn’t need fresh content. A search for “best smartphones 2026” does.
What this means for SEO
Update content when it genuinely becomes outdated — new data available, recommendations changed, industry shifted. Don’t just change the byline date and republish. Google’s semantic date analysis can identify cosmetic updates versus substantive refreshes. When you do update, make sure the changes are meaningful enough to shift the content’s semantic dating signals.
TitleMatchScore and OriginalContentScore
Two additional scoring signals deserve attention.
TitleMatchScore measures how well a page’s title tag matches a user’s search query. The documents describe this as “a signal that tells how well titles are matching user queries.” Importantly, the documentation suggests Google processes the full text of title tags regardless of display truncation — even if Google shortens your title in search results, the algorithm reads the entire thing.
OriginalContentScore evaluates content originality, with particular application to shorter content. The scoring range is 0-512 for short content. This contradicts the common assumption that short content automatically equals “thin” content in Google’s eyes. A 300-word article that provides genuinely original analysis may score higher than a 3,000-word article that rehashes widely available information.
A related attribute, contentEffort, appears to use large language models to estimate the effort required to create an article. This could help Google differentiate between content that required significant human expertise and content that was quickly assembled by scraping, translating, or algorithmically remixing existing material.
What this means for SEO
Title tags still matter — a lot. Write titles that accurately reflect the page’s content and match the queries you’re targeting. Don’t write clickbait titles that misrepresent the page. For content length, quality and originality matter more than word count. A concise, original piece can outperform a lengthy derivative one.
Topical Authority: SiteFocusScore, SiteRadius, and Embeddings
The leaked documents confirm Google uses vector embedding technology — specifically site2Vec — to create mathematical representations of websites and individual pages. These embeddings capture semantic meaning by converting content into multidimensional numerical coordinates.
Two metrics stand out:
SiteFocusScore measures how concentrated a website’s content is around specific topics. Higher focus scores reward sites with clear thematic identity. A site that publishes exclusively about commercial espresso machines builds a stronger topical signal than a site that covers espresso machines, cryptocurrency, and pet care.
SiteRadius measures how far individual page embeddings drift from the overall site embedding. Pages that stray too far from your site’s topical center may receive weaker ranking support. Google creates a topical identity for your website and assesses each new page against that identity.
Together, these signals explain why niche sites often outrank larger generalist sites for specific queries. A site with a tight SiteFocusScore and small SiteRadius in a particular topic area sends a stronger topical authority signal than a broad site that touches the topic occasionally.
The documents also suggest Google recognizes “topical borders” and evaluates whether content creates legitimate bridges between related subjects. A coffee equipment site expanding into coffee bean sourcing makes topical sense. The same site suddenly publishing content about home insurance does not.
What this means for SEO
Build topical depth before topical breadth. Establish authority in your core subject area with comprehensive, interconnected content before expanding into adjacent topics. When you do expand, create logical bridges — content that connects your core expertise to the new subject through a clear thematic relationship. Random topic expansion dilutes your SiteFocusScore.
Author and Entity Recognition
The documents confirm Google stores author information associated with content and can identify whether an entity is the author of a document. This is the clearest structural evidence for E-E-A-T’s “Experience” and “Expertise” components in the ranking system.
Google’s entity recognition goes beyond just reading bylines. The system can identify entities on a webpage and sort, rank, and filter them. For author entities specifically, Google appears to build profiles that accumulate signals across multiple pieces of content and multiple websites.
What this means for SEO
Author identity matters for ranking. Ensure your content has clear, consistent author attribution. Build author entities through profiles on authoritative platforms (LinkedIn, industry publications, speaking engagements), consistent use of author schema markup, and a body of published work that demonstrates genuine expertise in your subject area.
The smallPersonalSite Attribute
One of the more intriguing discoveries is an attribute called smallPersonalSite. The leaked documents don’t explain how it’s used. Mike King noted that Google could assign a ranking adjustment (“twiddler”) to this attribute that could either promote or demote small personal sites.
Given that Google’s Helpful Content Update and subsequent core updates in 2023-2024 significantly reduced visibility for many small niche sites and personal blogs, some analysts speculate the attribute may currently be used for demotion rather than promotion. This remains speculative — the function is genuinely unknown.
What this means for SEO
If you run a small personal site or niche blog, the existence of this attribute means Google is actively classifying your site type. Focus on demonstrating genuine expertise, building real audience engagement, and accumulating quality backlinks from established sources. The sites that survived the HCU updates typically had strong E-E-A-T signals, real author identities, and genuine user engagement — not just SEO-optimized content.
Demotion Signals: How Google Penalizes
Beyond positive ranking signals, the documents catalog several explicit demotion mechanisms.
anchorMismatchDemotion — triggered when inbound links’ anchor text doesn’t match the destination page’s content. Google evaluates both sides of a link for relevance. If your page about “hiking boots” receives dozens of backlinks with anchor text about “online casino,” the mismatch triggers demotion.
GibberishScore — detects low-quality, artificially generated content. The documentation specifically mentions content created through “low-cost untrained labor, scraping content and modifying and splicing it randomly, and translating from a different language.” Language models and query-stuffing detection help identify unnatural content patterns.
IsAnchorBayesSpam — a flag that identifies spam anchor text patterns. This works alongside anchorMismatchDemotion to evaluate your overall backlink profile for manipulative patterns.
Nav Demotion — penalizes sites with poor navigational experience. Confusing navigation leads to higher bounce rates and makes it harder for Google to crawl and index content effectively. Pages buried deep within submenus may not be indexed at all.
Panda — the documents reveal Panda uses embeddings for content quality assessment and references patents about modifying content rankings based on the number of independent links and NavBoost reference queries. Panda appears to function as a site-level quality modifier rather than a page-level penalty.
Sensitive topic whitelists — the documents reference whitelists for particularly sensitive query categories including elections, COVID-19, and travel. Content in these categories appears to be subject to additional quality filtering.
What this means for SEO
Every demotion signal reinforces the same theme: Google has detection systems for manipulative tactics, and those systems are far more sophisticated than most practitioners assumed. Anchor text manipulation, content spinning, artificial link building, and poor UX all have specific demotion mechanisms. Recovery from gibberish or spam penalties takes 3-6 months or longer, if recovery is even possible.
Link Signals: More Complex Than Expected
The leak confirms that backlinks remain a significant ranking factor — despite Google’s occasional suggestions that links are becoming less important. The documents reveal specific attributes the system evaluates:
Link source quality — the authority and trustworthiness of the linking website. Not all links are equal. A link from an established industry publication carries more weight than a link from a random blog.
Link source diversity — the number of unique referring domains matters more than the total number of links. A hundred links from one site are less valuable than ten links from ten different relevant sites.
Link relevance — Google evaluates whether the linking page is topically relevant to the target page. A link from a running gear review site to your running shoe product page carries more topical relevance than a link from an unrelated site.
Anchor text signals — the text used in the link provides topical context about the target page. But the anchorMismatchDemotion and IsAnchorBayesSpam signals mean over-optimized or mismatched anchor text can trigger penalties.
Link freshness — the documents suggest newer links may carry different weight than older ones. A continuous stream of new links from authoritative sources signals ongoing relevance, while a link profile that stopped growing years ago may signal declining authority.
The documentation also confirms Google uses multiple variants of PageRank — not the single metric that was deprecated from public view. Internal fields include toolbarPageRank, pageRank2, and other variants, each apparently serving different functions within the ranking system.
What this means for SEO
Link building isn’t dead — it’s the specific tactics that have changed. Focus on earning links through genuinely valuable content, digital PR, and authentic industry relationships. Prioritize diversity of referring domains over volume. Ensure anchor text is natural and varied. Avoid any scheme that produces a large volume of links from irrelevant sources with identical anchor text.
Twiddlers: Post-Algorithm Ranking Adjustments
The documents reference “twiddlers” — ranking adjustment functions that run after the primary scoring algorithm (Ascorer) within Google’s Mustang ranking system. Twiddlers can promote or demote pages based on specific signals, effectively fine-tuning the initial ranking output.
NavBoost itself appears to function partly as a twiddler — adjusting rankings after the initial content-and-link-based scoring based on real user engagement data. Other twiddlers may handle specific demotion signals (spam detection, quality filtering) or promotion signals (freshness boosting for QDF queries).
This layered architecture means Google’s ranking isn’t a single formula. It’s a pipeline where content passes through multiple evaluation stages, each of which can modify the page’s final position. Understanding this helps explain why pages sometimes rank well initially and then drop (or vice versa) — different twiddlers may activate as more data accumulates.
Rand Fishkin’s Big Takeaway: Build a Brand
After analyzing the full scope of the leak, Fishkin distilled his strategic advice into a single recommendation that has since become the most widely cited takeaway from the entire event:
“If there was one universal piece of advice I had for marketers seeking to broadly improve their organic search rankings and traffic, it would be: Build a notable, popular, well-recognized brand in your space, outside of Google search.”
This conclusion follows logically from the leaked signals. SiteAuthority rewards established, trusted domains. NavBoost rewards sites that generate consistent positive user engagement. Chrome data validates genuine audience behavior. Author entity recognition rewards known, credible experts. SmallPersonalSite classification distinguishes established brands from anonymous niche sites.
Every signal in the leak tilts the playing field toward brands that people know, seek out, and engage with — not just websites that are technically optimized for search crawlers. This doesn’t mean technical SEO is irrelevant. It means technical SEO is necessary but not sufficient. The sites winning in the long run are the ones building real audience relationships, brand recognition, and trust — signals that are hard to fake and that compound over time.
Practical Implications: What to Do With This Information
Prioritize genuine user satisfaction
NavBoost and Chrome data make user engagement a direct ranking input. Every page should be designed to thoroughly answer the query that brought the user there. Reduce bounce rates not through tricks (scroll hijacking, exit-intent popups) but through content that genuinely meets user needs. Track engagement metrics in your own analytics and treat declining engagement as an early warning for potential ranking drops.
Build topical authority systematically
SiteFocusScore and SiteRadius reward depth over breadth. Create comprehensive content clusters around your core topics. Avoid publishing unrelated content that dilutes your topical identity. When expanding into new topics, build bridges through content that logically connects the new subject to your established expertise.
Invest in brand building outside of Google
Fishkin’s advice isn’t abstract — it’s supported by the leaked signals. PR, speaking engagements, industry awards, podcast appearances, social media presence, and community participation all generate the brand recognition signals that feed siteAuthority, entity recognition, and the trust metrics that differentiate established brands from anonymous websites.
Maintain content freshness where it matters
Check whether your target queries trigger QDF (trending topics, recurring events, rapidly changing fields). For those queries, regular content updates with genuinely new information are essential. For evergreen topics, update when the content actually becomes outdated — not on an arbitrary schedule.
Build a natural, diverse link profile
Earn links from multiple relevant domains through valuable content, original data, and authentic industry relationships. Monitor your anchor text distribution for unnatural patterns. Avoid link schemes that produce volume without relevance or diversity.
Establish author entities
Create consistent author profiles across your website, social platforms, and industry publications. Use Person schema markup. Build a verifiable body of published work that demonstrates expertise in your subject area. Google’s entity recognition systems are watching.
Accept the sandbox period for new sites
New domains will face reduced visibility regardless of optimization quality. Plan for a 6-18 month ramp period. Focus early efforts on building brand signals, earning initial authoritative links, and creating a foundation of quality content rather than expecting immediate ranking results.
Monitor for demotion signals
Audit your backlink profile for anchor text mismatches and spam patterns. Review your content for gibberish scores — particularly if you’ve used translation, content spinning, or low-quality outsourced writing. Ensure your site navigation is clear and that important pages aren’t buried more than three clicks from the homepage.
What the Leak Doesn’t Tell Us
The limitations of this leak are as important as its revelations. The documents don’t reveal the weights assigned to each signal — we don’t know whether NavBoost is 10x more important than siteAuthority or vice versa. They don’t confirm which attributes are currently active versus deprecated. They don’t explain how signals interact — whether certain signals can override others, or whether there are threshold effects where a signal only matters once it exceeds a certain level.
Google’s caveat that the documents may be “outdated” is plausible — the most recent internal date reference suggests the documentation was current as of approximately August 2023. Google’s systems evolve continuously, and specific attribute names, weights, or implementations may have changed since then.
What the documents do confirm is the scope and direction of Google’s evaluation system. The fundamental signals — site authority, user engagement, content quality, link analysis, topical authority, entity recognition — are structural to how Google assesses the web. Even if specific implementations change, these categories of evaluation are unlikely to disappear.
Frequently Asked Questions
Does the leak prove that Google lied about its ranking algorithm?
The documents reveal clear contradictions between Google’s public statements and its internal systems — particularly regarding siteAuthority, click data usage, Chrome data integration, and the sandbox. Whether this constitutes “lying” or “strategic omission” is a matter of perspective. Google’s stated reason for not confirming these signals — preventing manipulation — is understandable from their position. But for SEO practitioners who made strategic decisions based on Google’s public guidance, the contradictions are significant.
Are the leaked documents still relevant in 2026?
The core signal categories — authority, engagement, freshness, topical relevance, link quality — are structural to how search ranking works. Specific attribute names or weights may have evolved since the documents were current (approximately August 2023), but the fundamental architecture Google uses to evaluate websites is unlikely to have changed at its foundation. The independent confirmation of NavBoost through the antitrust trial strengthens the documents’ ongoing relevance.
Should I change my SEO strategy based on the leak?
If your strategy was already focused on creating genuinely useful content, building topical authority, earning natural backlinks, and providing good user experiences, the leak validates your approach. If your strategy relied on tactics that Google publicly said were fine but the leak reveals are monitored for manipulation (aggressive anchor text optimization, click manipulation, cosmetic date updates), you should adjust.
How does siteAuthority differ from Moz’s Domain Authority?
Moz’s Domain Authority is a third-party metric calculated using Moz’s own methodology. Google’s siteAuthority is an internal metric calculated by Google’s own systems. The inputs, calculation methods, and weights are different. They’re conceptually related — both attempt to measure a domain’s overall strength — but they’re not the same metric and shouldn’t be treated as interchangeable.
Is Google’s sandbox a permanent barrier for new sites?
No. The hostAge attribute indicates differential treatment for new domains, not permanent exclusion. The sandbox appears to function as a probationary period where Google requires additional evidence of legitimacy before granting full ranking access. Building genuine brand signals, earning authoritative links, and publishing quality content during this period helps establish the trust signals needed to exit sandbox treatment. Practitioners estimate this period typically lasts 6-18 months for competitive queries.
What was the most important finding from the leak?
Industry consensus centers on NavBoost and the confirmation that user click data is a core ranking signal. This single finding reframes how SEO practitioners should think about optimization — not just making pages that search engines can crawl, but making pages that users genuinely want to engage with. The antitrust trial’s independent confirmation of NavBoost’s importance makes it the highest-confidence finding from the entire leak.




-150x150.png)

