August 10, 2026
On August 5, NPR's On Point devoted 38 minutes to a familiar apparel question: How can the same person wear an XS in one brand and an XL in another?
The program opened with Caroline Locke in Broomfield, Colorado, reading labels from her closet: small, medium, XS, large, 6, 10, 32, and a men's XL sweatshirt. Lululemon placed her at 12 to 14, or XL; Costco's Kirkland brand at a medium, or size 6.
The labels contradicted one another. The wardrobe still belonged to one person.
For shoppers, that inconsistency is frustrating and personal. For retailers, it creates a commercial problem: the product page cannot translate the shopper's experience into an answer for the item in front of her.
Retailers do not have to standardize the industry to solve that problem. They have to help one shopper choose one size in one product.
Calls for standardization are understandable. The history shows why one label has never been enough.
U.S. sizing research in the late 1930s and early 1940s measured roughly 15,000 women, but the sample did not represent a broad enough cross-section of the population. A federal commercial standard for women's sizing followed in 1958, was revised into a voluntary product standard in 1970, and was withdrawn in 1983.
The obstacle was not simply a lack of industry will. A size label is being asked to compress too much information: body dimensions, proportions, brand fit models, garment construction, fabric behavior, design intent, and personal preference.
Different brands also serve different customers. That variation can be useful. If every brand fit the same reference body in the same way, shoppers whose proportions differed from that body could have fewer options, not better ones.
The problem is not merely that a medium means different things at different brands. It is that the shopper is expected to determine which medium applies to her in a specific item.
Four layers of variation compound before a shopper reaches the size selector.
A waist measurement does not reliably predict hip measurement, thigh volume, rise, or torso proportion. Height does not map cleanly to size. Even apparently straightforward men's waist and inseam labels vary in practice.
Each brand chooses a fit model, target customer, grading system, and amount of ease. A tailored brand and a relaxed lifestyle brand can use the same label for intentionally different fits. Brands without strong customer data may also benchmark competitors, which can compound inconsistency.
Stretch, cut, construction, fabric recovery, and intended silhouette all matter. An oversized knit tee and a structured woven blouse can share a nominal size and fit nothing alike. Footwear varies across lasts, widths, materials, and use cases for the same reason.
Compressed development cycles can limit wear testing across a full size range and reduce the time available to assess shrinkage, twisting, recovery, and consistency. A shopper cannot always tell whether a surprising fit reflects deliberate design intent or product variance.
A universal label cannot resolve all four layers. Item-level guidance can translate them.
The On Point segment described three common responses to fit uncertainty. Each maps to a retail cost.
In a store, an experienced associate can absorb some of that uncertainty. Online, the shopper is often left with a static chart and product copy. PacSun found the same gap in its own customer data: the most common question shoppers asked its chatbot was how to understand the retailer's size chart across brands.
Support questions are the visible edge of the problem. The shoppers who never ask, then leave or order multiple sizes, are harder to see.
A closet is not automatically a clean dataset. Some garments are gifts. Some are tolerated rather than loved. Some no longer fit.
The useful signal is not mere ownership. It is the record of what a shopper bought, returned, exchanged, and kept.
Those outcomes can show how a shopper maps across brands, categories, products, and fit preferences. They are more informative than a click, a browse, or a self-reported answer because they add a feedback loop: the system can learn what worked and what did not.
That is the difference between intent data and outcome data. Intent data captures what a shopper viewed, searched, added to a cart, or entered into a quiz. Outcome data captures what happened after the choice.
True Fit is built on two decades of purchase and return outcomes across more than 100 million registered shoppers, 60 million unique products, 91,000 apparel and footwear brands, and more than $616 billion in transaction value. At the product page, True Fit uses that history to translate what is known about the shopper and the product into a size recommendation.
Not another chart. Not a generic average. An answer for the decision in front of the shopper.
True Fit and its retail partners report the same broad pattern across several implementations: when fit uncertainty is addressed before the order, conversion can rise while size-related return behavior falls.
These are retailer-specific case results, not neutral benchmarks or universal guarantees. Outcomes vary by catalog, category, integration, traffic, measurement method, and shopper adoption. The useful question is not whether every retailer will reproduce the same percentage. It is whether conversion and returns improve together.
Clothing sizes are inconsistent because each brand defines its own fit, and each product adds its own construction, fabric, and design intent. A universal label would not eliminate those differences.
The solvable problem is communication at the moment of choice. As fit expert Alison Hoenes told On Point, "the communication, the presentation to the customer of their fit and sizing" is where the industry has room to improve.
Shoppers do not need every brand to mean the same thing by "medium." They need a trustworthy answer about which size to order in the product they are viewing now.
Retailers that can answer that question with real outcome data have a chance to convert more of the demand they already have and prevent avoidable returns before they ship.
There is no enforced universal standard. Each brand defines sizes around its target customer, fit model, grading approach, and design intent. Fabric, cut, construction, and trend can also change how two products with the same label fit, even within one brand.
Not completely. Standard measurements could make labels easier to compare, but they cannot capture every combination of proportion, preference, fabric behavior, and product intent. The more useful intervention is to translate brand-and-item variation into a shopper-specific recommendation.
A size chart provides measurements or label definitions. It still asks the shopper to know her measurements, understand how the brand runs, account for the product's intended fit, and make the final inference alone. A personalized recommendation answers the actual question: Which size should I order in this item?
Size bracketing is ordering the same item in multiple sizes with the intention of returning the sizes that do not fit. Beyond the refund, it creates outbound and reverse-logistics expense, inspection and processing labor, inventory delay, and markdown risk.
True Fit's software learns from purchase and return outcomes across brands, products, and shoppers. It maps the shopper's fit history and preferences to the characteristics of a specific product, then delivers a size recommendation at the point of choice.
Yes. Footwear has the same cross-brand inconsistency plus variation in lasts, widths, materials, and intended use. ASICS reported a 150% increase in product-page-to-cart conversion after implementing True Fit for footwear.
Romney Evans is Chief Marketing Officer at True Fit, where he leads go-to-market strategy for fit intelligence across apparel and footwear.
Want to see what fit uncertainty is costing your catalog? Request a demo to see how True Fit maps purchase and return outcomes to your products and shopper base.