POV - How does the SCRUNCH acquisition matters the most at this point from Sitecore?

July 16, 2026 ·

POV - How does the SCRUNCH acquisition matters the most at this point from Sitecore?

Hello Everyone,


By now, everyone in the Sitecore ecosystem has heard it, Sitecore acquired Scrunch. LinkedIn is full of takes on what it means for the product roadmap, the partner network, the AEO/GEO category. Before I add to that pile, I want to talk about something which raised my eyebrows, and i want yours to raise too

Do you know? Search engines never had a traffic drop in last 20 years, but now first time in history, Search engines are seeing traffic drop.


We are witnessing the shift in a technology and use of it and we are at the very intersection of that change where your site audiences are changing, first time in a history people are not going to search engines, They are going to LLMs for their discovery / research / or day to day tasks.

Back in February 2024, Gartner published a prediction that quietly rattled the marketing world: by 2026, traditional search engine volume would drop 25%, with search marketing losing ground to AI chatbots and virtual agents, Generative AI was becoming a substitute for the search query itself, not just a new feature bolted onto it.

25% is not a small number. A 25% contraction in the channel that most B2B and B2C marketing strategies were built around, in a window of roughly two years.


Cloudflare CEO Matthew Prince expected the crossover by the end of 2027. It arrived in June 2026.

According to Cloudflare Radar, which monitors traffic across roughly 20% of the web, bots now generate about 57% of requests to web pages, compared with 43% from humans.

That means we have more bots than a human on the net already.




What if you have the best site and digital experiences and product catalog, but no one is visiting your site? the content which you curated in years of efforts and created those connected journeys are suddenly not showing results? 

What if the full marketing digital ecosystem feels like to rewritten for this new ERA of AI? 

That sounds scary right? If you are a CMO or a CTO or a Consultant or a Content creator, This article is for all of us who are in the same boat of this era of AI and how the audiences is no more humans, but bots.

You can only fix for which you know "Why" something is not working, Right now, When your website or experiences or campaigns are live, you know bots will visit, but if majority of your traffic comes from there, Then you do not know how to overcome it, Your content is not working, Not because its bad but for because it was created for HUMANS, now the audience is changed, you have to now think how can it be cited by bots or LLMs instead of the old age approaches which were only targeting humans.

We've lived through this kind of shift before

Think back to the dot-com boom. The internet flooded with websites and experiences, all built for one audience: humans, clicking, scrolling, reading. That was the paradigm for thirty years.

Now ask the uncomfortable question: what happens if your experience — no matter how brilliant the content, how polished the design, how strong the brand — simply stops getting that traffic? Not because it's bad. Because the audience reading it has quietly changed from people to machines acting on people's behalf.

That's not a marginal shift. That's the kind of shift that means your entire marketing ecosystem — built for human eyes over the last two decades — may need to be rewritten from the content layer up.

Old friends are getting new neighbors

We all used robots.txt from years, It has quietly governed how search engines crawl the web for about thirty years. In September 2024, AI researcher Jeremy Howard proposed a new file to sit alongside it: llms.txt. Its entire purpose is different from robots.txt, It doesn't tell bots where they can't go, it tells them where the content that actually matters lives, stripped of navigation, ads, and JavaScript that AI models don't need and can't easily parse. 

SEO too, Nearly every marketer has lived and breathed this discipline for years, now has new siblings: AEO, Answer Engine Optimization, and GEO, Generative Engine Optimization. Different mechanics, different signals, same underlying question: will your brand be the one an AI system cites when someone asks it a question your competitors also want to answer?

You can love this shift or resist it. You cannot ignore it. We're sitting at the exact intersection of the change. The real question isn't whether it's happening, it's whether you're ready for it.

Why the Scrunch acquisition is the important part of this story

This is exactly the gap Sitecore just spent real money to close. Sitecore's CEO Eric Stine framed the rationale directly that buyers are now consulting large language models ChatGPT, Claude, Gemini about products and vendors before they ever reach a company's website, and by the time they arrive, if they arrive at all, their opinion is often already formed.

Scrunch's co-founder Chris Andrew has described the company's founding bet, made back in 2023, in a single sentence worth sitting with: that the most important visitor to a website would eventually not be a person, but an AI agent acting on that person's behalf. Two years later, that bet looks less like a gamble and more like a forecast that arrived early.

This is why the acquisition matters to every Sitecore partner, customer, and brand they represent, Not because it adds another logo to the marketplace, but because it's aimed squarely at helping brands show up and get cited in an era where ChatGPT and Claude are becoming primary visitors to their sites, not secondary crawlers.

What Scrunch actually does, and why the name fits

Here's the simplest way I can describe it: Scrunch, quite literally, scrunches your website or a page.



credit : https://scrunch.com/blog/ai-site-crawlability-questions-answered

AI doesn't need the fancy design, the hero images, the animation, the layout choices your team debated for weeks. It needs the content that actually matters which can be clearly stated, easy to parse, free of the visual noise built for human persuasion. So Scrunch's Agent Experience Platform (AXP) creates an alternate, lightweight version of your page, Think of it as a "Human View" and an "AI View" running in parallel. It scrunches the page down, in the literal sense of the word, to just what a model needs to understand and cite it accurately, without touching the experience your human visitors actually see.

That distinction matters more than it sounds like on the surface. The compute cost or the "budget" a bot spends crawling and interpreting a page to process your full, heavy webpage versus your scrunched version is dramatically lower. Which means your value proposition, from an AI system's perspective, is dramatically higher. You're not just easier to find. You're cheaper for AI to read. And cheaper-to-read, well-structured content is exactly what gets selected and cited.

In a study with Akamai, pages enabled with Scrunch's AXP saw a 364% increase in brand presence for non-branded AI prompts and a 218% increase in citations across AI-generated results. In a separate case, Runpod used the platform to identify indexing and rendering issues limiting its AI discoverability, and reported a 400% increase in paying customers tied to that work.

The bottleneck every marketer already feels

Talk to almost any CMO or marketer right now and you'll hear the same thing, in slightly different words. They know AI is here. They know people are asking ChatGPT which product to buy, and that their competitors are the ones showing up in that answer. They know they should be optimizing for AEO and GEO.

What they don't have is the tool, the guide, or a clear picture of how to look at their own content the way an AI system does.

And that's the real bottleneck. If you don't know the intent of who or what is visiting your site, and you don't have a feedback loop telling you what content is actually working versus what's just sitting there unread, you can't make the decisions that matter like how to build your asset library, what content should power which campaign, how to build genuinely personalized journeys. It all comes back to one question, Who's visiting your site, and what do they actually want?

That's the gap Scrunch is built to close giving brands a map of who their visitors really are, human and machine, and the tooling to close the distance between the content they already have and what these next-generation AI engines need in order to cite them instead of the competitor.

Conclusion

This is the reason why it is one of the most important acquisition from Sitecore, You would never want your competitors to show up before you in LLMs, Instead you would want your partners and their billions of customers and their brand so they are cited on LLMs.

Your competitors are getting ready to be cited in LLMs, are you ready too? If not that's exactly the conversation worth having next.

Be cited. Not just searched ;-) Cheers !!!

Sitecore 10.2 - The Ghost 500 Error that Comes and Goes - How PrefetchData ends into Race Condition and Cache Corruption

June 29, 2026 ·

Sitecore 10.2 - The Ghost 500 Error that Comes and Goes - How PrefetchData ends into Race Condition and Cache Corruption

Hello Friends,

I believe this blog post is very important for everyone who is running Sitecore 10.2, because this is one of those issues which is very tricky to catch, very scary when you see it live, and very satisfying when you finally understand what is happening underneath. I will share my experience of what we faced, how we did a deep reverse engineering of Sitecore kernel, what we found and how we resolved it.

Issue we started facing

Our customer were sending new page publish in email and sms communication campaigns, and when someone clicks on that link it used to give 500 screen and on refresh it used to work, but first hit was giving 500 and user drop was happening, and because of this when we hit it, it always used to work so it was 

Our site started giving intermittent 500 errors with YSOD like following two exceptions

Exception 1: Index was outside the bounds of the array.

System.IndexOutOfRangeException: at System.Collections.Generic.List`1.Add (mscorlib, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b77a5c561934e089) at Sitecore.Data.DataProviders.PrefetchData.AddChildId (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.PrefetchData.AddChildrenForEmptyDefinition (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.Sql.SqlDataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.CompositeDataProvider+<DoGetChildIDs>d__94.MoveNext (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Common.EnumerableExtensions.ForEach (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.CompositeDataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.DataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataSource.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.GetChildrenCommand.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Engines.EngineCommand`2.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemProvider.GetChildren (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemProvider.GetChildren (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.ResolvePath (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.ResolvePath (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Engines.EngineCommand`2.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemProvider.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.ItemProvider.GetItem.GetLanguageFallbackItem.Process (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at n/a (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.CorePipeline.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.DefaultItemManager.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.DefaultItemManager.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemManager.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Pipelines.Response.GetPageItem.GetPageItemProcessor.GetItem (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Pipelines.Response.GetPageItem.GetFromRouteUrl.Process (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at n/a (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.CorePipeline.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.DefaultCorePipelineManager.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.DefaultCorePipelineManager.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Pipelines.PipelineService.RunPipeline (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Pipelines.PipelineService.RunPipeline (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Presentation.PageContext.GetItem (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Presentation.PageContext.get_Item (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.ExperienceEditor.Pipelines.Request.RequestEnd.AddPageExtenders.Process (Sitecore.Mvc.ExperienceEditor, Version=10.0.0.0, Culture=neutral, PublicKeyToken=null) at n/a (Sitecore.Mvc.ExperienceEditor, Version=10.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.CorePipeline.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.DefaultCorePipelineManager.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.DefaultCorePipelineManager.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Pipelines.PipelineService.RunPipeline (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Mvc.Routing.RouteHttpHandler.EndProcessRequest (Sitecore.Mvc, Version=8.0.0.0, Culture=neutral, PublicKeyToken=null) at System.Web.HttpApplication+CallHandlerExecutionStep.InvokeEndHandler (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a) at System.Web.HttpApplication+CallHandlerExecutionStep.OnAsyncHandlerCompletion (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a)

Exception 2: Value cannot be null. Parameter name: key

System.ArgumentNullException: at System.Collections.Concurrent.ConcurrentDictionary`2.TryGetValue (mscorlib, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b77a5c561934e089) at Sitecore.Caching.Generics.Cache`1+InnerBox.DoGetEntry (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Caching.Generics.Cache`1.GetValue (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Caching.Generics.Cache`1.ContainsKey (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.Sql.SqlDataProvider.EnsureChildrenPrefetched (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.Sql.SqlDataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.CompositeDataProvider+<DoGetChildIDs>d__94.MoveNext (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Common.EnumerableExtensions.ForEach (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.CompositeDataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataProviders.DataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.DataSource.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.GetChildrenCommand.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Engines.EngineCommand`2.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemProvider.GetChildren (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemProvider.GetChildren (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.ResolvePath (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.ResolvePath (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Engines.EngineCommand`2.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemProvider.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.ItemProvider.GetItem.GetLanguageFallbackItem.Process (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at n/a (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Pipelines.CorePipeline.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.DefaultItemManager.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.DefaultItemManager.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Data.Managers.ItemManager.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Resources.Media.MediaRequest.GetMediaPath (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Resources.Media.MediaRequest.get_MediaUri (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Resources.Media.MediaProvider.ParseMediaRequest (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Resources.Media.MediaRequestHandler.GetMediaRequest (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Resources.Media.MediaRequestHandler.DoProcessRequest (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at Sitecore.Resources.Media.MediaRequestHandler.ProcessRequest (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null) at System.Web.HttpApplication+CallHandlerExecutionStep.System.Web.HttpApplication.IExecutionStep.Execute (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a) at System.Web.HttpApplication.ExecuteStepImpl (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a) at System.Web.HttpApplication.ExecuteStep (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a)

Side Effect

Because of this 500 error, our site pages were showing 500 custom error page intermittently and our users were impacted badly. The most frustrating part was — refresh the page and it works fine. This makes it very hard to reproduce and very easy for people to ignore, but believe me, do not ignore this one.

1. What is the issue, why it is intermittent and when does it come?

So first of all, let me explain what PrefetchData is and why it is there.

When Sitecore serves a page request, it needs to know the children of many items to resolve paths, render layouts, and process pipelines. Rather than hitting the SQL database for every single child lookup, Sitecore has a mechanism called PrefetchData which acts like a warm-up cache. It preloads child IDs of items into an in-memory collection so subsequent lookups are fast.

Now here is the interesting part — this prefetch cache gets populated on the very first request after an app pool start or recycle. And that is exactly where the problem lives.

Why is it intermittent? Because it only happens during that very small window of time when the app pool has just started and multiple concurrent requests are hitting the site at exactly the same moment. In that race window, multiple threads are trying to write child IDs into the same internal collection simultaneously.

When does it come?

  • After every app pool recycle (scheduled or memory-based)
  • After IIS restart
  • After deployment
  • First hit after idle timeout of app pool

Once that window passes and the cache is fully populated, subsequent requests read from a stable cache and everything works fine. That is why refresh fixes it — by the time you refresh, the cache is already built.

2. How we Reverse Engineered Sitecore.Kernel.dll and found the root cause

This is the most interesting part of this investigation.

We had two different DLLs — the original Sitecore 10.2 Kernel and the patched version from Sitecore's hotfix. We used dnSpy (a free open-source .NET decompiler) to decompile both and compare the exact code.

Tool Used: dnSpy You can download it from: https://github.com/dnSpy/dnSpy

Steps we followed:

  1. Open dnSpy
  2. File → Open → Load original Sitecore.Kernel.dll
  3. File → Open → Load hotfixed Sitecore.Kernel.dll
  4. Navigate to: Sitecore.Data.DataProviders → PrefetchData class
  5. Look at the AddChildId method

Original Code (Broken)

public virtual void AddChildId(ID childId) => this._childIds.Add(childId);

That single line. That is the culprit. _childIds is a plain List<ID> which is not thread-safe at all. When multiple threads call Add() simultaneously on a List<T>, you get exactly what we saw — IndexOutOfRangeException because the internal array of List<T> gets corrupted during concurrent resize operations, also see lines in RED mentioned in above exceptions (top of the page) pointing to exact same code where it fails.

Hotfixed Code (Fixed)

public virtual void AddChildId(ID childId) { lock (this._childIdsLocker) this._childIds.Add(childId); }

And somewhere in the class there is now:

private readonly object _childIdsLocker = new object();

This is the classic monitor lock pattern in C#. Only one thread can enter the lock block at a time. All other threads wait outside the door. So now when 10 concurrent requests come in during startup, they queue up neatly and add their child IDs one by one without corrupting each other.

How does Exception 2 relate to Exception 1?

This was a chained failure. Exception 1 was corrupting the prefetch data. Exception 2 (ArgumentNullException in ConcurrentDictionary.TryGetValue) was happening because EnsureChildrenPrefetched was trying to use the cache that had already been corrupted — a null ID was stored where a valid key was expected. Fix Exception 1 at source and Exception 2 disappears automatically. They are not two separate bugs, they are cause and effect.

3. Sitecore KB and Official Patch

Sitecore has acknowledged this as a known issue. The official KB article is:

Known Issues - Retrieving the child items of resource items is not thread-safe 

The patch is available on Sitecore's hotfix SharePoint portal under: Sitecore XP 10.2 → Sitecore 10.2.3 rev. 013888 PRE → Platform Patch

However, a word of caution here — the patch available on the KB link is quite a large upgrade patch and it can bring along many other changes which you may not want to introduce in a stable production CMS. Same experience we had with aliases pipeline blog post I shared earlier.

Our Recommendation is to go with sitecore patch instead of smaller footprint patch

We initially tried writing a lightweight custom patch DLL to override just the AddChildId method via a config-based type override a much smaller footprint than swapping the entire Kernel DLL. But reverse engineering further revealed PrefetchData is constructed directly in code with new PrefetchData(...), with no config-driven factory hook. This explains why Sitecore chose to ship a modified core Kernel DLL for this fix rather than expose a smaller, safer extensibility point and is a good reminder that not everything in Sitecore's architecture is config-patchable, even when classes are marked virtual.

Quick Temporary Fix while you prepare the proper patch

If you want to stop the bleeding immediately without any code change, By creating a patch which disables the prefetch cache config to omit this caching part totally.

This disables the prefetch cache entirely so the race never happens. There will be a slight cold-start slowdown after app pool recycles but no more 500s. Good enough to stabilize production while you work on the proper fix.

NOTE: It was our decision to only patch things which was broken to be in control, Official hot fix and number of DLLs and config which we wanted to avoid as our instance was otherwise stable only, If you think you can go ahead and install the full hotfix given on the link

Observation after fix

We observed the site for 24 hours after applying the thread-safety patch and there were zero 500 errors from this issue. Happy customer, stable site.

Hell of sitecore aliases pipeline breaking the site with 500 error

May 31, 2026 ·

Hell of sitecore aliases pipeline breaking the site with 500 error

Hello Friends,

I belive this blog post is very important for everyone because, It has some very serious effect on working of your headless website, i will share my experience what we faced and how we resolved it

Issue we started facing

Our site started giving "Key cannot be null or empty" with YSOD like following 



Side affect

Because of this 500 error, Our site pages were showing 500 custom error page intermittently and our MAU (Monthly Active User) drop rate increased.

Sitecore KB

There is already Sitecore KB article talking about this error but the patch which is provided on this link is confusing as well as very huge and it could bring other issues along with it as that upgrade patch also has lot of other things too which i did not want to introduce in our stable CMS.

Known Issues - Retrieving the child items of resource items is not thread-safe

Observation

Though the surfaced exception was looking similar and giving same error and behavior given on this article, We looked closely the inner exception and stack trace where we noticed following in bold

System.ArgumentNullException:   at System.Collections.Concurrent.ConcurrentDictionary`2.TryGetValue (mscorlib, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b77a5c561934e089)   at Sitecore.Caching.Generics.Cache`1+InnerBox.DoGetEntry (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Caching.Generics.Cache`1.GetValue (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Caching.Generics.Cache`1.ContainsKey (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.DataProviders.Sql.SqlDataProvider.EnsureChildrenPrefetched (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.DataProviders.Sql.SqlDataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.DataProviders.CompositeDataProvider+<DoGetChildIDs>d__94.MoveNext (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Common.EnumerableExtensions.ForEach (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.DataProviders.CompositeDataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.DataProviders.DataProvider.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.DataSource.GetChildIDs (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Nexus.Data.DataCommands.GetChildrenCommand.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.Engines.EngineCommand`2.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.Managers.ItemProvider.GetChildren (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.Managers.ItemProvider.GetChildren (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.ResolvePath (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.ResolvePath (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Nexus.Data.DataCommands.ResolvePathCommand.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.Engines.EngineCommand`2.Execute (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.Managers.ItemProvider.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.Managers.ItemManager.GetItem (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.AliasResolver.get_Item (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Data.AliasResolver.Exists (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Pipelines.HttpRequest.AliasResolver.Process (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at n/a (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Pipelines.CorePipeline.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Pipelines.DefaultCorePipelineManager.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Pipelines.DefaultCorePipelineManager.Run (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at Sitecore.Web.RequestEventsHandler.OnPostAuthenticateRequest (Sitecore.Kernel, Version=17.0.0.0, Culture=neutral, PublicKeyToken=null)   at System.Web.HttpApplication+SyncEventExecutionStep.System.Web.HttpApplication.IExecutionStep.Execute (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a)   at System.Web.HttpApplication.ExecuteStepImpl (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a)

   at System.Web.HttpApplication.ExecuteStep (System.Web, Version=4.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a) 

I knew the function of aliases resolver of Sitecore and why it exists, but did not know why that pipeline even executing if we do not have any aliases defined in Sitecore, So i was little surprised with this, Even if it is running, It should just exist because there are no aliases defined in sitecore.

So i did further digging to know what is happening behind the scene, Here are the findings 

1. What is the "AliasResolver"?

The AliasResolver is a standard processor located in the <httpRequestBegin> pipeline. Its primary job is to look at the incoming URL path and determine if it matches a predefined Sitecore Alias (configured under /sitecore/system/Aliases).

If a match is found, it maps that pretty/short URL to the actual content item path in the tree and sets it as Context.Item.


2. Why is it giving "Value Cannot Be Null (Parameter Name: Key)"?

The Layout Service Multi-threading: When your Headless/JSS application hammers the /sitecore/api/layout/render/jss endpoint, concurrent async requests cross paths in the ASP.NET Core  .NET  pipeline.

and our observation also reavelad that, this error is only coming when the request is of  "/sitecore/api/layout/render/jss"



3. We are close, but what is the issues and how to resolve it?

The Shared Resources Cache: To resolve aliases, the AliasResolver safely checks the cache or queries the child collection of the aliases root. In Sitecore 10.2, Sitecore moved several system items (including templates and system settings) into Read-Only Resource Files (.dat files on disk) to speed up performance.

The Dictionary Race Condition: When multiple concurrent Layout Service threads attempt to resolve items or read the children of these resource-backed elements at the exact same time, a race condition occurs within an internal collection (such as PrefetchData or Dictionary).

The Crash: One thread corrupts the internal array or returns a null value where a key string or ID was strictly expected. When the concurrent thread picks it up, the code drops a low-level .NET ArgumentNullException: Value cannot be null. Parameter name: key (or an IndexOutOfRangeException), bubbles up through the AliasResolver, and throws a 500 Internal Server Error.

Solution

There are three solutions to this issue, And each solution depends on what kind of issues you are running into and as dictionary race condition could come without aliases too, so you will need to observe your stack trace before you apply any of below 

When you decompile the kernel.dll, You will see that there is an setting item which decides, If this resolver should be executed or not


If you see, if that setting is TRUE then only it will try to go to Sitecore and find it and if its not found in Sitecore, it will try to find it in resource file and that is where it fails.

So, If you are using aliases in your application, You can create a custom AliasResolver processor that immediately aborts processing if the current request is directed at the Layout Service endpoint.

Approach - 1: Write a Custom Resolver

using Sitecore.Pipelines.HttpRequest;

using System;

namespace YourNamespace.Pipelines.HttpRequest

{

    public class CustomAliasResolver : AliasResolver

    {

        public override void Process(HttpRequestArgs args)

        {

            // Abort immediately if this is a JSS Layout Service call

            if (args.Url.FilePath.StartsWith("/sitecore/api/layout/render/jss", StringComparison.OrdinalIgnoreCase))

            {

                return;

            }

            // Otherwise, fall back to standard Sitecore Alias resolution

            base.Process(args);

        }

    }

}

And patch it in via Configuration

Replace the default AliasResolver with your newly optimized class:

<configuration xmlns:patch="http://www.sitecore.net/xmlconfig/">

  <sitecore>

    <pipelines>

      <httpRequestBegin>

        <processor type="Sitecore.Pipelines.HttpRequest.AliasResolver, Sitecore.Kernel">

          <patch:attribute name="type">YourNamespace.Pipelines.HttpRequest.CustomAliasResolver, YourAssemblyName</patch:attribute>

        </processor>

      </httpRequestBegin>

    </pipelines>

  </sitecore>

</configuration>

Approach - 2: Delete the AliasResolver pipeline completely using patch, if you are not using aliases or simply patch below setting item, By default the value is true, but make it false.


OR

<configuration xmlns:patch="http://www.sitecore.net/xmlconfig/">

  <sitecore>

    <pipelines>

      <httpRequestBegin>

        <processor type="Sitecore.Pipelines.HttpRequest.AliasResolver, Sitecore.Kernel">

          <patch:delete />

        </processor>

      </httpRequestBegin>

    </pipelines>

  </sitecore>

</configuration>

Approach - 3: Upgrade to Sitecore newer version or patch

Our scenario was different as we were getting clear alias pipeline stack trace, but if you observe same error either in content management instance or on content delivery with stack trace given on below link, Please update to the patch given in Sitecore KB below 

https://support.sitecore.com/kb?id=kb_article_view&sysparm_article=KB1001823

I observed the site for 24 hours after this patch, and no 500 errors, and a happy customer and we observed that drop rate was decreased and site started functioning normally and MAU increased.

BTW - i have raised the feature requests about changing the pipeline so that it should only execute code for resource files and overwhelm the race condition if aliases are used, if they are not used, it should just work without the upgrade or patch.

Why SitecoreAI - Getting into the shoes of the customer how to select right CMS

April 10, 2026 ·

Why SitecoreAI - Getting into the shoes of the customer how to select right CMS

Hi Team,

Lately, I have been talking to lot of our customers / potential customers and having pre-sales demos where one question always comes is "Why Sitecore"

Now this question can be for any product which is out for sell.

And as a technician I always get into product technical features, but at the same time as a pre-sales guy, it also makes me think, surely all competitive products have same features, so definitely answer to this is not in the technicalities. 

If you step back and think, we are also a customer in our daily life and buy lot of things, what is that process we go through? When we buy, how can your customer decide if this is a right fit for you or not, why we select A over B? Is it price? is it service? Is it a brand? Is it about features? Is it about brand loyalty? 

When it is a technical product, I am sure it cannot start with the technicalities of the product or selecting product itself, 100% not, I feel decision is always business strategy first and then going for a selection of the right tool

If we talk in context of selecting CMS, we cannot start with the available CMSs out there and selecting one of it first and then think about fitting the business strategy or problem into it, It will backfire 100%

So, How to select right CMS?



Right path is first strategizing the digital transformation strategies by asking questions like

  • What is the end goal?
  • How can I give connected experience to my customer?
  • How can we automate things?
  • How easy it is for a business to go to marker fast?
  • What is the marketing team is trying to do?
  • Are we getting the right data, for right audience? and have right campaigns? 
  • Is it future ready?

And lot more questions like these should be discussed first before any RFP (Request For Proposal) is out or you are getting into any discussions with any vendor

So crux is,

A customer can decide whether a CMS/DXP is the right fit by starting with business strategy and experience goals (not the tech), translating those into measurable requirements, and then evaluating platforms against a small set of “non‑negotiables” (security/compliance, performance, integration, and operability) plus “future-ready” capabilities (composability, omnichannel, scalability, and upgrade burden). The decision is typically made by a customer's DX team combined, so the evaluation must produce a clear business case with quantified impact, risk reduction, and total cost of ownership (including upgrade/maintenance effort).


Key Findings of how the decision should be made by a customer team

  • Fit is determined by experience outcomes first: omnichannel cohesion, content discoverability, and connected customer journeys should be defined and measured before selecting technology.
  • Buying is team decision-driven: success requires mapping stakeholder priorities (Marketing, IT, Security, Compliance, Product, Support, Finance) to a single decision scorecard and business case.
  • The strongest decision drivers are operability + risk + scale: security/compliance, performance/reliability, integration flexibility, and the ability to evolve without heavy upgrades typically outweigh feature checklists.

Detailed Analysis

1) Start with Strategy and then translate to Experience Requirements

See the image above of three steps of Strategy, Experience , Technology - Please observe technology is always last.

A. Strategy (why)

  • Growth goals: acquisition, conversion, retention, upsell
  • Operating model: centralized vs federated content teams, global/local governance
  • Time-to-market expectations: campaign velocity, experimentation cadence
  • Risk posture: regulatory exposure, availability requirements, incident tolerance

B. Experience (what)

Define the experiences you want to deliver across channels:

  • Omnichannel cohesion: same customer identity, consistent content and offers across web/mobile/email/in-store/portal/support
  • Content findability: search, taxonomy, personalization rules, content reuse, localization
  • Connected experience (CX in DX): journey continuity (handoffs across touchpoints), personalization, segmentation, integration with CRM/CDP/commerce


C. Technology (how)

Only after A and B, evaluate whether a platform can deliver those outcomes under your constraints:

  • Architecture fit: monolith vs headless vs composable/hybrid
  • Integration pattern: API-first, middleware compatibility
  • Deployment/operations: cloud, SaaS, self-hosted, DevOps maturity


2) Core “Fit” Questions (presented in customer's language)

Use these questions as the platform fit gate:


Omnichannel experience

  • Can the platform deliver content consistently to all current and planned channels (web, app, kiosk, partner portals, IoT, etc.)?
  • Can it support reusable content models (structured content, components) rather than page-only authoring?

Content is easy to find

  • Does it have strong content modeling, taxonomy, metadata, and search (both for editors and end users)?
  • Can teams govern and locate assets quickly (DAM integration, tagging, workflows)?

Current setup: "good" vs "future-ready"

  • Are you stuck spending most effort on upgrades and maintenance versus building new experiences?
  • Can you add new channels/features without replatforming?
  • Does the vendor roadmap align with your direction (AI tooling, composability, privacy changes, new channels)?

Ecosystem decision (not a single decision maker)

  • Can you create a business case that different stakeholders will accept (Marketing speed, IT operability, Security risk reduction, Compliance readiness, Finance TCO)?


3) Key driving decision makers (who influences what)

A CMS/DXP choice typically has these decision-makers and drivers, It depends a lot on communication from customer's transformation team and vendor team to talk and have the decision makers aligned on all areas to be able to decide which platform and which tech. they should go.


Generally, following are the key points which different stakeholder would like to have in any CMS.

1) Marketing / Digital Experience

  • Faster publishing, campaign velocity, personalization, experimentation
  • Editorial UX, workflow, approvals, localization

2) Product / Business Owners

  • Conversion rate, customer satisfaction, retention
  •  Ability to launch new journeys/features quickly

3) IT / Architecture

  • Integration, scalability, maintainability, DevOps fit
  • API-first, modularity, observability, incident management

4) Security

  • Identity/access control, secrets management, audit logging
  • Vulnerability management, patching model, data protection

5) Compliance / Legal (Healthcare, Automotive, E-commerce implications)

  • Regulatory alignment (HIPAA/PHI handling if applicable; PCI for payments; privacy laws)
  • Data residency, retention, audit trails, accessibility requirements

6) Finance / Procurement for TCO calculations

  • Total cost of ownership (licenses + hosting + implementation + ongoing ops)
  • Vendor risk, contract terms, predictable costs

7) Customer Support / Operations

  • Reduced content errors, fewer incidents, easier rollbacks
  • Consistency across self-service and assisted channels

4) Decision Drivers (what typically “wins”)

Now that we have all of it covered, But customer team's influence will based majority on below points which could be key decision drivers for them.

 1) Performance & Reliability

  • SLA/uptime, latency, global delivery, caching/CDN strategy
  • Peak traffic handling (campaign spikes, seasonal events)
  • Monitoring/observability (logs, metrics, tracing)

2) Security

  • SSO (SAML/OIDC), RBAC/ABAC, MFA, least privilege
  • Secure SDLC, penetration testing, vulnerability disclosure
  • Encryption (at rest/in transit), key management, audit logs

3) Compliance

This is one of the important points, if there are domain specific requirements, how much compliant the CMS is, even if it does everything above but its not compliant, The decision will be not to use this until it becomes compliant.

  • Industry-specific: HIPAA (health), PCI (payments), privacy regulations, accessibility (WCAG), records retention
  • Data residency and vendor subprocessors

4) Future-ready (scale + adaptability + ease of use)

  • Composable capabilities: ability to swap components (search, personalization, commerce) without replatforming
  • Content reusability across channels
  • Developer productivity (SDKs, APIs, CI/CD)
  • Editor productivity (components, previews, approvals, localization)

5) “Upgrade effort” as a first-class KPI (how to reduce wasted effort)

Many organizations pick a platform that becomes an "upgrade loop." A fit check should explicitly score:

  • Upgrade frequency and effort: hours per month/quarter spent on patching, regressions, re-testing
  • Breaking-change risk: how often upgrades require rework
  • Operational burden: hosting, scaling, security patching responsibilities

What reduces upgrade effort:

  • SaaS-managed platform where patches and infra are handled by vendor (with clear release management)
  • Strong backward compatibility and tooling (migration tools, deprecation policies)
  • Automated testing and staging environments that mirror production

6) How to build the business case for a combined team decision

A good business case ties platform capabilities to measurable outcomes:

1) Value (revenue / growth)

  • Faster launches → more campaigns → lift in conversion/lead volume
  • Personalization/targeting → higher AOV/retention

2) Cost savings (operational)

  • Reduced engineering time on upgrades/maintenance
  • Fewer incidents and lower support burden
  • Reduced vendor sprawl (consolidation)

3) Risk reduction

  • Compliance exposure reduction (auditability, access controls)
  • Security posture improvement (patching, logging, incident response)

7) Practical evaluation scorecard (use this to decide “right fit”)

Customer team (shown in above wheel diagram) should now put high level points on the board and start the evaluation process based on different BUs, Practical way would be to have a score card for every points and have the evaluations done in different categories, Some could be like below

  1. Experience enablement -  (omnichannel, personalization, preview, workflows)
  2. Content operations - (modeling, reuse, localization, governance, search)
  3. Integration & architecture - (API-first, commerce/CRM/CDP, middleware)
  4. Security & compliance - (SSO, audit logs, certifications, data controls)
  5. Performance & reliability - (SLA, global delivery, DR, observability)
  6. Operability & upgrade burden - (patching model, release cadence, migration effort)
  7. Total cost of ownership - (license + implementation + run)
  8. Vendor viability & roadmap - (support model, community, product direction)
  9. Team fit - (editor usability + developer productivity)

NOTE: Here you can evaluate CMS tool as well as different vendor too in the same score card.

8) "Latest offering" and "tech stack used" (how customer should interpret these)

Instead of "latest features" in isolation, customer should ask:

  • Which features directly improve our KPIs (speed-to-market, conversion, operational cost, risk)?
  • Are “latest offerings” stable, supported, and used by reference customers at our scale?
  • Does the tech stack align with our talent pool and operating model (SaaS vs self-managed; frameworks, APIs, integration tooling)?
  • As AI is everywhere, One of the criteria is definitely AI capabilities, How much of AI features are backed into CMS, how can it empower the CMS user or content authors.  

Summary

The whole exercises which i gave above is to help customer decide when they are going under digital transformation and want to select CMS from pool of different CMS available

I have had many experiences where if above exercises are not done, It creates friction among customer and vendor team, because customer is expecting things which the selected tool or vendor can not do

So, i really hope this extensive guide will all the customers out there and use it as a starting point 

NOTE: All of the above images are generated with google gemini.


Zero to Hero - A real life RCA of exact issue in Sitecore Managed Cloud environment

November 27, 2025 ·

Zero to Hero - A real life RCA of exact issue in Sitecore Managed Cloud environment






Hello All,

The purpose of today's post is to share a real life burning and escalated scenario which was new to me and how did I approach it and how big the escalations were and what was the outcome

Sitecore's goodwill was at stack not because Sitecore is not capable of handling it but just because our environment was Sitecore Managed Cloud, and any issue that comes if its infra, back end code, front end code will be first pointed as Sitecore issue and that is where our consultancy and experience will play a role to prove that it is not Sitecore issue. 

Issue we faced

Out of the blue our site started giving "504 Gateway Time-out", and it was reported that almost everyone is getting this error, but when we used to browse the site, everything looked good and never 504.


504 Gateway Time-out error tells that, That the request went to Content Delivery servers of Sitecore from gateway, but gateway did not get response in time from those CDs and hence it gave time out error.

One thing we knew was, there is something wrong with backend, but backend can be Sitecore, databases, Solr, Azure itself, Backend can be anything which is involved in processing the requests/response, so we knew our scope is very broader, felt like "Finding a needle from the haystack"

But the side effect of this was customer was loosing on MAU (Monthly Active Users) on which their business model is working and it is the most important KPI for them.

 

Troubleshooting without initial clue


We had to start somewhere, because we had no clue why for certain user this issue is coming, so first thing that we do is to check Sitecore logs to see what happened, because this happened for other users and we were just informed.

Just to mention, we are on Sitecore 10.2 XM on manage cloud set up.

1) Sitecore Logs


We could not find any errors in Sitecore logs which could lead to anything like this, there was no specific pin pointing happening from the logs.

(Later on, we found out that because Sitecore never gave the response back to gateway in order so that something can be logged in the logs, because gateway timed out before the response.)

2) Azure application insights


Because we did not find anything in Sitecore logs, we moved to application insights and used transaction search, where you can just type any string and it will search in logs, and it gave us something

What we found were lot of SOLREXCEPTIONs in the time period where this 504 was reported, from that we found that there is an issue with SOLR whenever 504 happens, so we went to solr and checked the logs 

SOLR was choking when these 504 scenarios were happening




We had little clue now what is happening but then we investigated further graphs of response time and server errors during that time of SOLR and here is what we found

Avg. Response Time




Clearly, there was so many exceptions which we were seeing in LOGs were also observed in this graph too.

But our question was, Though SOLR has less power but why it works almost all the time and what happens in the specific scenario that it chocks up? 

So, we set up some theory around the scenario after getting on a call with marketing team and understanding the scenarios of what exactly they are doing and happening.

More from application insights

We observed that in specific time duration when this happens, All the request/response takes huge time to load and we tried to see in performance tab of AVG. time a request is taking in that specific time duration and what it reveled was also something pointing to SOLR

We observed that some of the layout services queries were taking huge time to get back the response, and all of our components read data from SOLR, so there is possibility that if SOLR chokes, it will take time to respond back.


Now if any request is taking more than 1 min. the gateway is bound to time out and give 504 Gateway-timeout error.

Now, we had to find what is causing this, Because if we hit the same page, it works just fine and in that time frame request finishes in perfect time, So we had to drill down to find if it's really a code issue or a caching issue where requests are being made to server all the time and timing out? 

4) Sitecore Ticket

For safer side we also created a SC ticket with these dumps, in initial findings it was also showing long running queries and choking queries of layout services, and sometimes it showed one component being slow and sometimes it showed other components being slow, which was quite confusing

and memory and CPU analysis kept sending us to BOX-1 of slow running layout queries which were timing out the request, but now we knew that these could be side effect of the main cause of some kind of automatic traffic hitting the site which will choke any server.

5) Memory Dump & CPU analysis

We also enabled auto heal and memory dump and site slow down reports etc. now we had memory dumps collected, but all the dumps were showing layout service choking which we already knew.

6) Rendering caching & Performance enhancement & Cache tuning

Because the reports were also showing that layout service is taking time, we had to dig into every rendering, review implementation of third-party calls, APIM calls, DB audit etc. to make sure nothing is miss configured and which could lead to this

Because this was already escalated and P1 issue, if we do this audit and performance enhancement, it will take time, so we decided to remove some of the rendering to see what happens, even after removal of some of the renderings which were depending on third party API, issue still came very next evening

Now, we had to find a way to first know the scenario of what exactly is happening, so we decided to talk to the team who were reporting this issue and we found out that they are observing whenever they are sending out marketing communication with website link, at that time only this issue is occurring

We even increased DefaultHTMLCache and other cache values as we had enough memory to cache things, but no luck with that.

We knew that those many users are not clicking the link but still our site is showing lot of traffic, so something is happening when SMS are sent, and it is not Sitecore or a backend issue from configuration or code perspective.

So, we looked at it from a fresh perspective and set up our hypothesis. 

Setting up hypothesis around the inputs from marketing team.


We now set the theory, of this is happening when there are lot SMS being sent, because marketing team is sending lot of WhatsApp and SMS comms and it seemed that exactly at the same time site is getting outages

All of the timings matched too, whenever site gave 504, all of the time marketing team sent communications

So, we had now something to think about, because on the website we already did the load testing and server scaling etc. was in place, we were sure we had a good set up which could handle the traffic.

7) Load Testing

Though we decided to do a load testing again, and our load testing was showing all results ok with the same URLs which marketing team was sending in whatsapp or SMS marketing.

So again, we were sure that, it's not about the legitimate traffic but it is something to do with SMS and without clicking also burst of traffic is coming like DDoS

8) Server scaling

We decided to put more power on CDs, so we introduced 4 additional instances just to observe for couple of days so we can take out some reports and observe what is happening.

Even if we did that, we observed same amount of 504 request, and we observed following which was pointing to the same scenario of traffic in few seconds.


 

9) Azure Front Door logs

One more reporting and azure diagnostic query we run to check what is coming on AFD and what is giving 504, we took that report out in excel and something we observed 

We observed that, there were some IP ranges which were hitting the sites in few seconds, so if you observe the above graph too

Within seconds site got 4-5 thousand hits, this behavior is not of users, there will be delay of few seconds, here are the IPs which we found 

64.233.173.0/24

66.102.6.0/24

66.102.7.0/24

66.249.82.0/24

66.249.83.0/24

66.249.84.0/24

66.249.88.0/24

192.178.11.0/24

74.125.215.0/24


Majority of them are google bots, but question was why those many requests are coming? and this was happening only when communications were being sent directly to user's mobile as SMS.

10) Further research yielded interesting facts


We went back to our SMS sending partner who sends out bulk SMS and asked them for a report in those dates of outages, and they also sent out circular that, they are seeing sudden high traffic spikes exactly at the same time when SMS are sent, and those can't be users.

Our research on AI platforms and google search also revealed this behavior of pre-fetching of the link to the site for some security and safe browsing link scanners services reasons, and also, we got this from our partners 

"Recent updates in Google Messages introduced enhanced security scanning and automatic link previews. These previews are generated by background preview agents that fetch the URL in advance to show a link preview to users.

Since these fetches imitate real user behaviour, they were being counted as actual clicks, resulting in inflated click numbers and sudden traffic spikes on your website.
 
Google uses a set of proxy/crawler IPs (including the ranges you shared) to pre-fetch URLs found in SMS messages. Because of this, you may observe:
Multiple requests to your URLs within 2–3 seconds
IPs from Google subnets appearing as traffic sources
Sudden server load spikes right after SMS campaigns
“Bot clicks” recorded even before users interact"

Wow, we were thrilled to learn this behavior, we could understand google bot hitting sites but not aware about this auto clicking of URLs imitating user click behavior, where traffic looks legitimate 

and the hypothesis we established seemed to have a proven theory ahead of us.

Solution


As we know that actual cause was something else and layout choking, SOLR choking and whole infrastructure getting choked was just side effect of these BOT auto clicks and sending huge traffic.

1) We blocked following user agents along with above IP ranges on WAF.

·       Mediapartners-Google

·       Googlebot

·       AdsBot-Google

·       Applebot

·       HeadlessChrome

·       YandexBot

·       facebookexternalhit

·       facebookcatalog

·       WhatsApp

·       YandexImageResizer

·       YandexMobileBot

·       YandexImages

·       YandexAccessibilityBot

·       YandexRenderResourcesBot

·       YandexUserproxy

·       GoogleMessages

·       okhttp

·       /web/snippet/

·       http4s-blaze

·       SkypeUriPreview Preview/0.5

·       Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/139.0.0.0 Safari/537.36

It has been a week, and everything seems to be working fine, and we can see on AFD reports that those IPs are now getting blocked and now all the legitimate requests are turning into 200 status code.

We double checked this behavior with Sitecore and they gave following links, Which is already offered by Sitecore Managed Cloud for DDoS kind of attack or traffic.

How-to's - Sitecore Managed Cloud – DDoS attack mitigation steps

Support Information - Sitecore Managed Cloud Standard (MCS) PaaS 1.0 — DDoS IP Protection

Next action item is to put a rate limit rule based on the campaign URLs to make sure site does not get overwhelmed as these IP ranges may change and user agent could change too. 

Summary Points

Because when you are working as a Sitecore partner, for customer it's a Sitecore platform, and goodwill of Sitecore goes into stake, in this specific scenario same happened, just because it is on Sitecore Managed Cloud.

But finally, we were able to pull that off and customer was convinced that it is not Sitecore issue :) 

Most important things we learn are from actual pressure situations, and backing ourselves in our intuitions which comes from experience, especially when you have not come across situation like these, here are some take away I will always call out below important points when approaching these kinds of situations 

1) Holding your ground

2) Believe in your hypothesis

3) Never give up 

4) Use your intuitions which comes from experiences 

5) Find alternatives 

6) Start with fresh ideas

7) Take a break to get more fresh ideas

8) Think out of the box

9) Work as a team, to have more brain working

10) Don't leave any stone unturned 

Thank you Manglesh Vyas for being there when needed the AFD reports or azure graphs etc. and my backend mate Kiran Sawant who always made sure that if we want to send any change of backend be it caching or fine-tuning resolvers was sent and checked during this P1 issue.


 


Sitecore  - How to show a new marketing promotional page on the same URL as existing home page

October 27, 2025 ·

Sitecore - How to show a new marketing promotional page on the same URL as existing home page

Hi Team,

Today i will share one of the solution that we did for one of our customer, I am sure you will or you already might have came across such requirements and found your self in multiple option/solutions and trying to find best suited one for your customer, here is the story and solutions we thought of and finally selecting one out of it which was the best in all scenarios

Also the solution was required in time sensitive deadline before their social marketing campaign begins so we had to come up with the solution and implement and go live before it.

Customer Requirement

They were doing a brand refresh, so whole site supposed to be revamped, With new user interface and UX, but that is a longer route, by the time we create that fully new site for them, they wanted to have a teaser home page, or a new home page to be shown just to give the visitor a feel of what is coming and they can market it using social campaigns.

So their need was, Whenever users visit a website (www.blahblah.com for example), Instead of old home page, it should show a new home page with new logo, new menus and new everything, and from there clicking on a "Go to website" button they can go to old website.

Looks fairly simple task right? But now it came with following caveats 

1) URL should not be changed, so no physical routes, but it should still load the new home page, Still preserving the old home page as we had to send user to old home page from that new marketing page.

2) Once user click on the "Go to website" from new marketing page, User should go to old home page without changing the URL again, and after that user should be in the old pages only, all the navigation should work "as-is"

3) Whenever they hit the site in the browser, always it should first load the new teaser (new home page)

4) From old site logo and all the links which takes user to a home page, should load the old home page only, because now user is into the old site.

5) From email campaign links or social media links or any links via search engine which points to website's home page, should first load the new teaser page (home page).

6) Site is multilingual, so new marketing page also had that language switcher and once language is changed the same marketing page should remain.

Possible solutions with their limitations

Because there are multiple ways of doing one thing, and we started thinking about possible approaches to this requirement, and we also noted down the drawback that came with each approach.

1) Change the startItem to new page

One of the first approach was to change the startItem of site from Sitecore site settings, But because this approach was every time loading the new home page, and to load the old home, page we needed to use the url like www.blahblah.com/home as a physical route, because now Sitecore is not treating this old home as start item, to access this item we had to give the physical path in the URL, this looked ok initially, that on the "Go to home" CTA on a new teaser page we can add www.blahblah.com/home route and it will load the old home page, But because this startItem is being referenced on Sitecore's link generator, all the links of the site needed to be changed or custom URL resolver needed to be written which can transform www.blahblah.com/home/about-us to simply have www.blahblah.com/about-us, and this approach we put on the shelve. 

2) Use redirects from old to new and new to old

Our customer's infra team said, They will use redirects from Azure Front Door to decided when to load which page, But this approach we clearly cancelled with reasoning that, we clearly need some identifiers to decide when to load which page and it can not be done without changing lot more from infra side, so it was also put on shelve.

3)  URL Referrer approach 

Our internal links when clicked will have referrers, our JS can identify if there is referrer and if its our internal page then allow them to go to old home page, and when there is no referrer or a refer which is not the current site then send them to new teaser page , but flow here was it will have flicker because JS can only be executed once page is loaded and it will not have the user experience we wanted.

4) UTM query string approach 

It had flow same as point-3, but then we also thought about reading these on our resolvers and then remove the query string before sending user to a page, otherwise all pages will have those UTM SOURCE attached, Also it might end up in too many redirect if not handled properly, Clearly we thought there can be better approach.

5) Use Personalization 

We were looking for a fast and efficient way to achieve this behavior, Though it can be achieved with personalization but lot of existing components needed to be changed and we wanted to stay away from changing existing components and let them work and behave as-is.

After all these thought process, we clearly knew, we need two things

1) Some how we need to know the way to send user to old and new page using some mechanism.

2) To do point-1 we needed identifier or a decision maker when to do that redirect

Final & Working Solution

ItemResolver with referrer as the identifier

 public class HomeItemResolvercs : HttpRequestProcessor

    {
        public override void Process(HttpRequestArgs args)
        {
            string referrerUrl = System.Web.HttpContext.Current.Request.UrlReferrer?.AbsoluteUri;
			// this is my site name in IIS, if you have it differently hosted, just change this to yours
            string mySiteUrl = "homepagepoc.localhost"; 
            // When path is of root and there is no referrer that means its organically hitting the site home page
            if(string.IsNullOrEmpty(referrerUrl) && System.Web.HttpContext.Current.Request.FilePath.Equals("/"))
            {
                var teaserItem = GetTeaserItem();
                if (teaserItem != null)
                    Context.Item = teaserItem;
            }
           // when path is of root and when there is a referrer that means, home link was clicked from somewhere 
           else if(!string.IsNullOrEmpty(referrerUrl) && System.Web.HttpContext.Current.Request.FilePath.Equals("/"))
            {
              // if from search engine or email campaigns, if referrer comes then also it should load teaser page, for our internal pages, it should always open old home page
if(!referrerUrl.Contains("localhost") && !referrerUrl.Contains(mySiteUrl)) { var teaserItem = GetTeaserItem(); if (teaserItem != null) Context.Item = teaserItem; } // for all other links load home page var homeItem = GetHomeItem(); if (homeItem != null) Context.Item = homeItem; } } private Item GetTeaserItem() { var sitecoreDB = Sitecore.Context.Database; var sitecoreQuery = $"/sitecore/content/Home/newhome"; if (sitecoreDB != null) { var teaserItem = sitecoreDB.SelectSingleItem(sitecoreQuery); return teaserItem; } return null; } private Item GetHomeItem() { var sitecoreDB = Sitecore.Context.Database; var sitecoreQuery = $"/sitecore/content/Home"; if (sitecoreDB != null) { var homeItem = sitecoreDB.SelectSingleItem(sitecoreQuery); return homeItem; } return null; } }


Small Explanation

Basically, The custom resolver has two main condition and checks 

1) It only executes if the request is for home page, So that other requested routes of other pages will load "as-is" without any issue

2) Because the code only executes when request is made for a home page, now the decision to load either teaser page (new home page) or old home depends on the referrer, 

3) From internal site links, if you want to open the new home page from any link, for example press release or something, you can simply add rel="noreferrer" (we used this for language switcher so that it loads the marketing page only but with new selected language)  to your anchor tags and it will load the teaser page, so you are in control.

When the page is directly hit in a browser, it does not have a referrer, and when someone clicks internal link or "Go to home" link on the page, It will automatically attach the referrer and that is our identifier to decide if we should load the new home page or the old home page, See the demo in action below

So, all organic clicks from search engine/ emails / direct browser url will always load the new teaser page and else all other links of the site, internal links will always load the old home page

Cheers !!! Small but useful solution :) 

NOTE: Don't forget to add the processor for this custom pipeline :), to help you I have put together the code and processor on the GITHUB link   


PS: My colleague Nelson   also worked on couple of other approaches and did POC on his end which reads the custom query string and when user is redirected to the page, those query string will be removed.