It Looks Like an Attack
WAF false positives, part 2 of 3: the managed rules teams switch off, and the tracking data a WAF mistakes for attacks.
Third-party scripts put data into your requests that managed rules read as attacks. The usual fix, switching the rule off everywhere, throws away the protection along with the false positive. This part covers which managed rules teams switch off most, then the second story: a marketing-pixel cookie that a WAF reads as SQL injection, and a tracking parameter blocked as a file extension. Plus a detour into what counts as a “real user” when so many legit clients aren’t humans in browsers.
In Part 1, a fitness app’s WAF was blocking 4.5% of the traffic to its own Share button. When a rule causes trouble like that, some teams choose to disable it entirely, and that happens a lot.
Which managed rules teams switch off most
Some rules cause constant friction with legitimate traffic (human or automated), while others are perfectly safe: a browser-fingerprint rule, in a world where browsers change shape every release, is trouble; a geo rule usually isn’t. Here’s how often each managed rule, across the configurations we see, ends up switched from block to count:
Two of the rules teams disable most turn up in this series: SizeRestrictions_BODY is the rule that ate Trailhound’s workouts in Part 1, and the anonymous-IP rules come up at the end of this part. And the usual “fix” is to switch off the whole rule on every path, right after it broke something visible, instead of adding a scoped exclusion for the one affected endpoint. The protection is then gone everywhere, including where it was doing its job. Some overrides are deliberate and sound. Many others are rules that broke something once, got switched off everywhere, and stayed off.1
Disabling a rule only helps once you know it’s the problem. In the next case, that was hard to find out.
Case Files
Stories from real WAF logs, involving the same managed-rule families you’re probably running. Story 1 can be found here. Story 2 is below, and Story 3, telling the tales of "bad languages", is in a future post in this series.
It looks like an attack:
a marketing cookie flagged as SQL injection
Roamio sells experiences: kayak tours, cooking classes, day passes to places with roller coasters. The page that matters most is checkout, and everything in this story happens on the way to it.
How anomaly scoring works
Roamio’s WAF runs on Azure Front Door with Microsoft’s managed ruleset, which descends from the OWASP Core Rule Set and, like it, works on anomaly scoring: rules don’t block, they add points. Each SQL-like or XSS-like match adds points by severity, and once the total crosses a threshold the request is blocked. It’s a sensible design: a single minor match can’t block a request on its own.2 The downside is that several harmless matches can add up to a block.
A marketing-pixel cookie that looks like SQL injection
Then the team added a marketing pixel, a small script that reports ad conversions back to an ad platform. From that day on, every visitor’s browser picked up a cookie named _mkt_sid, whether or not they’d ever seen one of the ads, and sent it with every request to the site: browsing, search, checkout. The value is a machine-generated ID like this:
1741862930514::hV7tQw2Lx--Kp4_9-zR8.2.1741862930733.0In the middle is --, the SQL comment marker from every ' OR 1=1 -- you’ve ever seen. Rule Microsoft_DefaultRuleSet-2.1-SQLI-942100 matches the double dash and flags the request as SQL injection, even though there’s no SQL anywhere in it.
The cookie ID is random. Most users’ _mkt_sid never happens to contain --. But sometimes the generator produces a double dash, and that user is now flagged as an SQL injection attempt on every request they make, all the way to checkout. In one month: about 85,000 blocked requests from this one cookie.
(Component)
Why the cookie bug can’t be reproduced
This one is hard to debug because it’s random. A developer or QA engineer opens the site, gets a clean cookie, and everything works. Blocked users who clear their cookies get a new ID, and the problem vanishes. Meanwhile the logs fill with “SQL injection attempts” from ordinary people trying to book a cooking class.
None of the offending code was written by Roamio. The value came from an ad platform, and the WAF that blocked it came from yet another vendor.
The fix: a scoped cookie exclusion
(Component)
A tracking parameter blocked as a restricted file extension
Like most companies its size, Roamio runs more than one WAF. AWS WAF sits in front of other services, including one that receives Roamio’s first-party analytics (telemetry sent to its own domain, to get past ad-blockers). Those beacons carry a Google-Tag-Manager-style payload in the query string, ending like this:
/experiences/lisbon-food-tour/similar?…&data=event=gtag.config
AWS WAF’s RestrictedExtensions_QUERYARGUMENTS (part of AWSManagedRulesCommonRuleSet, meant to block dangerous file types like .bak, .config, .ini) scans the query arguments, finds .config, and blocks. The “file” is a tracking parameter whose value happens to end in .config, one of the extensions on the rule’s list. Unlike the cookie, this isn’t random: the beacon always ends in gtag.config, so every user on that flow gets blocked, every time.
(Component)
Aside: what counts as a “real user”?
Trailhound (Part 1) and Roamio were content mistakes. Another family of rules ignores the content and judges the sender: IP, network, client type. Geo, hosting-range, anonymizer, and bot rules all rest on an assumption that stopped being true a while ago: that a real user is a human, in a browser, on home internet. Today a huge share of legit clients aren’t: partner integrations, mobile apps, server-to-server calls, AI agents acting for real people. AWS’s AWSManagedRulesAnonymousIpList is a good example: HostingProviderIPList blocks cloud/hosting ranges, which can be exactly right on a login form and wrong on the public API those same clients are supposed to call.
(Image)
Where geo, hosting-range and bot rules belong
Whether that block is a false positive is ultimately a business decision. Some parts of a site should be reachable only from residential ASNs; others have no business being restricted that way. The trap is that HostingProviderIPList rides in the same group right next to AnonymousIPList, the familiar “block Tor, proxies, and VPNs”, so it gets switched on without anyone consciously deciding to. (One customer even blocked bots from their signup page, a plain GET, turning away the crawlers and agents that bring signups.)
The hard part is knowing where a policy belongs: which endpoints need a geo, hosting-range, or bot restriction, and which are stuck behind one they never needed. Surfacing that is a good chunk of what we do. On telling helpful agents from hostile ones, our piece on agentic traffic covers this in detail.
Whether a rule judges the content, like Roamio’s cookie, or the sender, like a hosting-range list, the fix has the same shape: scope the rule to where it belongs instead of switching it off everywhere.
- Some overrides tell smaller stories: a curious number of teams single out
UserAgent_BadBots_HEADER, quite possibly clearing the way for their own DAST scanner, which announces itself honestly in its User-Agent. Understandable, but anyone can copy that User-Agent. Recognize your scanner by its source IP range or a pre-shared header, and let the bad-bots rule keep doing its job. - Severity matters more than count. A critical-severity match (an SQL-injection signature, say) is weighted heavily enough to cross the anomaly threshold by itself; that’s what happens below, one match, one block. Lower-severity matches usually have to stack up before they cross it.



