Shopify Indexing Issues: Cracking the 'Crawled, Currently Not Indexed' Mystery

Hey fellow store owners!

We’ve all been there: staring at our Google Search Console reports, seeing that dreaded “Crawled – currently not indexed” status for pages we know are important. It’s frustrating, right? Especially when you’ve put in the effort to make your store look great, only for Google to seemingly ignore some of your hard work.

Recently, a fantastic discussion unfolded in the Shopify Community, started by @TimberOps, who was facing this exact puzzle. They had product and collection pages that were structurally identical, with the same templates and metadata, yet some were indexed while others were stuck in Google limbo. What made it even more confusing was an intermittent issue with subordinate sitemaps timing out. This thread brought out some brilliant insights from experts like Roman79, IFGeCommerce, v.marychenka, and others, offering a clear path to diagnose and fix these tricky indexing anomalies. Let’s dive into what we learned.

Understanding "Crawled, Currently Not Indexed"

First off, it’s crucial to understand what “Crawled – currently not indexed” actually means. As Roman79 and IFGeCommerce pointed out, this isn’t a technical block. It means Googlebot successfully fetched your page, read its content, but then chose not to include it in its index. It’s a quality or relevance decision, not a discovery failure.

This is distinct from a sitemap timeout issue, where Google might not even reliably discover your pages in the first place. The community strongly advised treating these as two separate problems until you find concrete evidence linking them. Why? Because if Google already crawled the page, a sitemap timeout can’t be the direct reason it wasn't indexed. It might affect future recrawls or discovery of other pages, but not the initial indexing decision for an already-crawled URL.

Diving Deeper: The Content Quality Connection

So, if Google has seen your page but decided not to index it, what gives? Nine times out of ten, as Roman79 highlighted, it comes down to content quality and uniqueness. Google wants to provide users with the best, most relevant results. If your page looks too similar to others, or lacks substantial unique value, Google might decide it's not worth indexing.

Think about it: are your product descriptions unique? Are you using supplier-provided copy verbatim across many products? Are there many variants of the same item with very little distinguishing content? These are common culprits. Even if two pages share an "identical template," their content might not be "different enough" for Google’s high standards. As Maximus3 aptly put it, “Crawled not currently indexed is not a permanent rejection, it’s just a decision not to put it in the search index right now.”

Your Action Plan for Content:

  1. Compare Indexed vs. Non-Indexed: Francesco Guiducci from IFG eCommerce suggested building a “matched sample.” Take five indexed products and five non-indexed products from the same collection or using the same template.
  2. Analyze in Search Console: For each URL, use Google Search Console to compare:
    • The Google-selected canonical URL
    • The last crawl date
    • Whether crawling and indexing are allowed
    • Any sitemap or referring page information
    • The rendered content Google actually received
  3. Review Page Value: Go beyond technical structure. Does the non-indexed page offer substantially more unique information? Does it satisfy a clearer search intent? Is it too similar to other URLs in your catalog? Strengthen internal links to your priority products from high-traffic pages, as @ahsandoesntcare recommended.

Unpacking the Sitemap & Infrastructure Puzzle

While the sitemap issues might not directly cause "Crawled – currently not indexed" for pages already fetched, they are still critical for discovery and recrawling. @TimberOps initially found that their main sitemap was healthy, but subordinate sitemaps (like https://fitzgeraldcustomwoodcraft.com/sitemap_collections_1.xml?from=279639949415&to=298397466727) were intermittently failing or timing out for Googlebot. Mustafa_Ali and IFGeCommerce both emphasized the importance of pushing Shopify Support on this, as it points to a potential serving-layer or CDN issue.

A common mistake, highlighted by Maximus3, was testing a sitemap URL in the main Search Console URL Inspection tool. Remember, that tool is for pages, not sitemaps. For sitemaps, you need to go to the "Sitemaps" report in GSC to see its status, last read date, and any specific fetch errors. Even if a child sitemap times out now and then, as @dropfeed noted, Google doesn’t forget URLs it has *already* read from it.

Your Action Plan for Sitemaps:

  1. Keep Shopify Support Engaged: If you can reproduce intermittent timeouts on subordinate sitemaps, continue working with Shopify Support. Provide them with exact timestamps and the affected sitemap URLs whenever a failure occurs. This gives them concrete data to correlate with their server logs.
  2. Monitor GSC Sitemap Report: Regularly check the Sitemaps report in Google Search Console. Look for "Success" status and recent "Last Read" dates for your subordinate sitemaps. This report is your single source of truth for how Google is interacting with your sitemaps.

The Structured Data Surprise: JSON-LD Fix

Mid-discussion, @TimberOps discovered a new twist: Google Search Console reported an “Invalid top level element ‘string’” error in the JSON-LD output on some product pages. This was a fantastic catch, as malformed structured data can definitely confuse Google and impact indexing.

Community member v.marychenka quickly pinpointed the cause: the Product JSON-LD block was being output as a quoted string instead of a proper object. This often happens due to an extra | json filter in the theme code.

Step-by-Step JSON-LD Fix:

  1. Backup Your Theme: Crucial first step! Always duplicate your theme before making any code edits. Go to Online Store > Themes, find your live theme, click Actions > Duplicate.
  2. Edit Theme Code: Go to Online Store > Themes, click Actions > Edit code for your duplicated theme (or live theme if you’re brave and have a backup).
  3. Search for Structured Data: In the theme editor, use the search bar (usually top left) to search all files for structured_data.
  4. Locate and Remove the json Filter: Look for lines within an The | json part is likely the culprit. You want the output to be a direct JSON object, not a string representation of it. The fix is to ensure the JSON-LD schema is correctly formed, often by checking if an extra | json filter is applied where it shouldn't be. Remove just that | json part if it's causing the issue, ensuring the output within the script tag is valid JSON. This might involve looking at snippets like {% render 'product-structured-data' with product %} and then editing that snippet.
  5. Validate Fix in GSC: After making the edit, go to Google Search Console. In the left menu, check the "Merchant listings" and "Product snippets" reports. If they were showing the ‘string’ error, click "Validate fix" in each report. Google will recheck the listed pages, giving you a start date to track against your "Crawled, currently not indexed" pages.

Your Google Search Console Toolkit

Google Search Console is your best friend for these kinds of issues. Don't just glance at the overview; dig into the reports:

  • URL Inspection Tool: Use this for specific product or collection pages. It tells you the exact reason Google skipped a URL, shows the last crawl date, the Google-selected canonical, and how Google rendered the page. You can also manually request indexing for your most important pages here.
  • Sitemaps Report: As discussed, this is where you monitor the health of your sitemaps, checking for fetch errors and ensuring Google is reading them regularly.
  • Indexing > Pages Report: Get an overview of all your pages, their indexing status, and crucially, which canonical Google decided to use.
  • Merchant listings and Product snippets reports: Essential for identifying and validating fixes for structured data errors like the JSON-LD issue TimberOps found.

Ultimately, solving "Crawled – currently not indexed" often comes down to a two-pronged approach: optimizing your content for uniqueness and value, and systematically eliminating any technical barriers like structured data errors or sitemap delivery problems. The community thread showed us that while Shopify handles a lot of the technical heavy lifting, there are still specific areas where our attention to detail, especially with content and structured data, makes all the difference. Keep those Search Console reports open, experiment with your content, and don’t hesitate to lean on the community and Shopify Support for the infrastructure-level quirks.

Share:

Start with the tools

Explore migration tools

See options, compare methods, and pick the path that fits your store.

Explore migration tools