Split testing puts you ahead of almost every brand on Amazon. It still won't catch this.
If you’re already split testing, you’re ahead of almost everyone. Here’s the check even testers skip.
Last month, I went full nerd and did something I’d been wanting to do for years. I used Claude to analyze 900+ PickFu split testing polls (now 1038+) from 34K (now 36K) respondents across the 850+ eCommerce brands we’ve worked with. The result was a database that lets me look up trends, spot patterns, and answer almost any question about a client’s creative before we even start designing. This has reframed how my team and I make every creative decision and I’m sharing all of it here in Design Proof. Twice a week: data-backed creative insights to help you convert better on Amazon and eCommerce.
It’s time for another Tuesday split testing insight, but before I get there, I want to tell you about something I built over the weekend.
I had a deep session with the AI I lovingly refer to as Claudette (if you’ve been reading a while, you’ve heard me talk about her before). And yes, she’s a her, because she can multi-task ;-)
I got access to a much more powerful version of Claude, free within my usage plan, and I burned through that allowance fast. By the time I was done, I’d spent over $1,000 building two different tools for the business. I truly believe it was money well spent, down to the last penny.
I’ll tell you about one of them today and save the other for another post, but this one ties directly into everything this newsletter is built on. It takes our whole testing library and turns it into a “God view” dashboard for my team. Anyone on the team can pull up a poll, a category, a pattern, at any time. We can filter it, slice it, question it, and run our own designs against it. We’re using it to predict what we should be testing next, put together client reports with real data behind them instead of piecing it together by hand, and dig into our own designs against everything we’ve ever tested.
It also just quietly surfaces trends to the team, the kind of pattern one of us might notice and never think to mention to anyone else. Now it’s sitting in one place everyone can see.
I was in that dashboard this weekend when something caught my eye: gender divergence.
Quick context before I get into it, because I don’t want you to read this and think “sure, but do I even need to worry about this.” Most brands on Amazon never split test anything at all.
If you’re already running PickFu polls or Amazon experiments, you’re doing something most of your competitors aren’t doing at all. This post isn’t for the brands who haven’t started testing yet. It’s for the ones who have, because even inside testing, there’s a check I don’t think most people run, including plenty of brands who test constantly.
Let’s start here: A kids’ shampoo brand about to launch, ran a main image poll and got a winner. On the surface, it’s exactly the kind of result you close out and ship without a second thought.
Then we split the vote by gender. The winning image pulled 50% of male respondents. Among women, it pulled 24%. Kids’ personal care is a category where the primary buyer is female by a wide margin, a category-level estimate we use as a starting point when we don’t have the brand’s actual Seller Central data, not a hard number for this specific brand. Even using that as a rough guide, the image that won the poll was the image the minority of the panel actually preferred.
Nobody sets out to build a shampoo strategy around what men think of it. It happens quietly, buried in a table most people scroll past on the way to the headline number. With this particular client, this mattered because the we realized her packaging wasn’t resonating with the audience she needed to attract. She was able to rebrand before launching because of this catch, saving her from lost serp clicks.
But here’s where this actually matters even more, especially once you’re an 8 or 9-figure brand. At that size, your creative isn’t a rough sketch anymore, it’s built to speak to a specific audience. If your buyer skews toward one gender and you’re not testing against that specific audience, any new designs your launch to Amazon out your conversion at risk.
Compare that to a snack bar brand that ran a similar test. 74 respondents, a clear winner, 69% overall. When I split that one by gender, the winning image pulled 89% of women and 48% of men, a 41-point gap. For a snack bar brand with a predominantly female buyer, the overall winner and the right answer for the primary buyer were the same image. Same size gap, completely different story. That one was fine to ship as-is.
Two polls, two gender splits of similar size. The difference was never the size of the gap. It was whether the group actually driving the overall win was the group that buys the product. That is what your should be testing for.
I went back through our full library and pulled every poll with a gender divergence above 20 points. There are 85 of them, roughly 1 in 11 of the 932 tests we’ve run. Most confirm the obvious and point the same direction as the overall winner. But a meaningful chunk show the thing that should stop you: the overall winner was set by the segment least likely to actually buy the product.
The divergence isn’t random, either. It clusters in categories where the buyer skews hard one way: Pet Supplies (13 polls), Grocery (10), Supplements (10), Kitchen & Dining (8), Tools (8). Those are exactly the shelves where “who showed up in the panel” and “who actually buys” can quietly come apart.
Below I walk through six of these polls in detail, including a notebook brand where the same pattern showed up in the opposite direction (male respondents controlling the win in a category that skews heavily female), and a full product-image-stack test where the split was sharp enough to flip the winner entirely.
Here’s the mechanism to watch. A PickFu poll aggregates preference across every respondent. The overall winner is what the majority of whoever showed up preferred. That’s only your answer if the people who showed up are a reasonable stand-in for the people who actually buy. When gender preference diverges significantly and your buyer skews heavily toward one segment, the headline result can tell you exactly the wrong thing.
One honest caveat before you take this as gospel: this is an audit habit, not a law. A 30-person panel split by gender leaves you with maybe 15 people per segment, so a jump from 24% to 31% is really just two or three people changing their vote. I haven’t isolated every detail here, and I’m not claiming the exact percentage is precise down to the point. What I am confident in is the pattern: when the gap is wide and the segments disagree on the winner, it is worth two minutes to check who was actually driving it.
Most brands never check this. When the winner is obvious, the gender breakdown is the data section most people scroll past.
Pull up your most recent PickFu poll right now. If there’s a gender breakdown visible in the results page, check whether the overall winner also won with the gender that makes up the larger share of your buyers. If yes, you’re done. That result is clean. If the winner was set by the minority segment, that’s worth understanding before you ship.
The paid section covers exactly how to read it.
🔒 Now Let’s Get To The Good Stuff
Upgrading to paid is like getting the version of my brain that isn’t filtered for a general audience. Less than one Starbucks order per month, and you get the strategy I charge clients real money for — broken down so you can use it yourself.
Below the paywall:
A detailed breakdown of six of these high-divergence polls with what drove each split
Whether the poll type (main image, A+ content, or full image stack) changes the odds of this happening
How to find your buyer’s gender distribution in Seller Central - A two-step check that takes about two minutes and tells you whether your next poll result actually applies to your buyer.
Subscribe to unlock.



