T-shirt sizing estimates work using relative labels such as XS, S, M, L and XL instead of numbers. The team compares items against each other rather than estimating duration. It is fastest and most useful for large, poorly understood items where numeric precision would be false comfort.
Key Highlights
- T-shirt sizing uses relative labels rather than numbers, which stops teams debating false precision on work nobody understands yet.
- It suits epics, features and roadmap level items. For sprint level stories that are already well understood, story points usually serve better.
- Its main advantage is speed. A backlog of thirty items can be sized in under an hour, where numeric estimation would take considerably longer.
- Non technical stakeholders engage with sizes far more readily than with story points, which makes it useful in mixed groups.
- The most common mistake is mapping sizes to numbers, such as S equals 2 and M equals 5. That recreates story points with extra steps and loses the benefit.
- Sizes are only meaningful within one team. Comparing a Medium in one team against a Medium in another is not valid.
What T-Shirt Sizing Is
T-shirt sizing asks a simple question about each item: relative to the others on this list, is this small, medium or large?
The scale usually runs XS, S, M, L and XL, though teams vary it. Some drop XS, some add XXL for items that are clearly too big and need breaking down.
The technique is relative, not absolute. Nobody is asked how many days a Medium takes. The only judgement required is whether this item is larger or smaller than that one, which humans are considerably better at than absolute estimation.
That comparison is the whole mechanism. Ask a group how long a piece of work will take and you get a wide spread of confident, wrong answers. Ask which of two things is bigger and you get fast agreement.
T-shirt sizing sits alongside other agile estimation techniques rather than replacing them. Most teams use more than one, applying different techniques at different levels of detail.
If you are building these facilitation skills for a Scrum Master role, our free CSM practice test is a quick way to check your grounding in the framework first.
How to Run a T-Shirt Sizing Session
The mechanics are straightforward and the facilitation determines the quality.
Start with a reference item. Pick something the team knows well and agree it is a Medium. Everything else is judged against it. Without a reference, the first few items get sized arbitrarily and everything after inherits that error.
Take one item at a time and ask for a size, not a discussion. People commit to a size first, then explain. Discussing before committing lets the loudest voice anchor everyone else.
Reveal simultaneously. Whether that is holding up cards, typing in a chat at a count of three, or a digital tool. Sequential answers produce herding, where the second person adjusts toward the first.
Discuss only the outliers. If most of the team says Medium and one person says XL, that gap is the valuable part of the session. The XL person has usually spotted something the others have not, or has misunderstood the item. Either way it needs surfacing.
Re-size after the discussion. Once the disagreement is explored, ask again. Do not average the answers, and do not let the most senior person simply decide.
Timebox the whole session. Twenty to thirty items in an hour is a reasonable pace. Sessions that run long usually mean the items are not understood well enough to size, which is a refinement problem rather than an estimation one.
The output is a sized backlog. The conversations that happened along the way are the more valuable product, because they surface assumptions that would otherwise appear mid Sprint.
When to Use T-Shirt Sizing
The technique fits some situations well and others badly, and matching it correctly matters more than the technique itself.
| Situation | Good fit | Why |
| Epics and features | Yes | Too large and uncertain for meaningful numbers |
| Roadmap and quarterly planning | Yes | Rough relative sizing is all the decision needs |
| Initial backlog triage | Yes | Fast enough to size a large backlog quickly |
| Mixed technical and business groups | Yes | Non technical stakeholders engage with sizes readily |
| Sprint level user stories | Usually not | Items are understood well enough for finer granularity |
| Sprint capacity forecasting | No | Sizes do not aggregate into a reliable forecast |
| Comparing across teams | No | Sizes are only meaningful within one team |
The pattern is that t-shirt sizing suits coarse decisions about large things, and struggles where teams need arithmetic.
That distinction is easy to state and often ignored. A team sizing its sprint backlog in t-shirts, then trying to work out how much fits in the Sprint, has chosen a technique that cannot answer the question being asked of it.
T-Shirt Sizing Compared With Story Points
These two get positioned as alternatives when they mostly operate at different levels.
| T-shirt sizing | Story points | |
| Scale | XS to XL labels | Modified Fibonacci numbers |
| Granularity | Coarse | Finer |
| Speed | Very fast | Slower |
| Best level | Epics, features, roadmap | Sprint ready stories |
| Aggregates into velocity | No | Yes |
| Non technical accessibility | High | Lower |
| Risk of false precision | Low | Higher |
Story points support velocitybecause they are numeric and additive. That is their main advantage and the reason teams use them for Sprint level work.
T-shirt sizing deliberately gives that up. Sizes do not add, so there is no velocity, no burndown and no forecast. What you get in exchange is speed and the absence of false precision.
The honest framing is that they solve different problems. Story points answer how much can we take into this Sprint. T-shirt sizing answers which of these twelve initiatives is bigger than the others, which is a question that arises months before anything reaches a Sprint.
Many teams use both. Epics get t-shirt sized during roadmap discussions, then broken down and pointed once they are close to being worked on. That progression matches how understanding actually develops.
The Conversion Trap
This is the mistake that quietly removes the benefit, and it is extremely common.
A team adopts t-shirt sizing, then someone asks how to forecast with it. The natural answer seems to be a mapping: XS equals 1, S equals 2, M equals 3, L equals 5, XL equals 8. Now the sizes add up and produce a number.
At that point the team has story points with a longer route to them. Every advantage of t-shirt sizing has been discarded.
The conversion also introduces a genuine error. The whole reason to use sizes on large uncertain items is that precision is unavailable. Converting an XL into an 8 asserts that it is exactly four times an S, which nobody actually believes. The number arrives with an authority the underlying judgement does not support, and then gets used in planning as though it were measured.
A related version is teams that record both a size and a point value for the same item. That is not more information, it is the same judgement expressed twice, and the numeric one will win any argument because it looks more rigorous.
If a team genuinely needs numbers, use story points. If precision is unavailable, use sizes and accept that forecasting will be rough. Choosing sizes and then reverse engineering numbers gives the disadvantages of both.
What Makes an Item Each Size
Teams work faster once they have shared reference points. These are common patterns, though every team should define its own.
XS. A trivial change with no unknowns. A configuration update, a copy change, a small fix in a familiar area. Often not worth sizing separately.
S. Understood work in a familiar area, no dependencies, done something similar before. Comfortably fits inside a Sprint with room to spare.
M. The reference point. Meaningful work with some complexity, still well understood, fits inside a Sprint. Most items should cluster here if the backlog is well refined.
L. Substantial. Likely to consume most of a Sprint, may carry dependencies or unfamiliar areas. Worth asking whether it can be split.
XL. Too large to work on as one item. Not really an estimate but a signal that this needs breaking down before it can be planned.
Treating XL as a flag rather than a size is one of the more useful habits available. When an item comes back XL, the productive next question is what smaller piece of this would deliver value on its own.
A backlog where most items are L or XL is not a backlog ready for planning. It is a list of intentions, and refinementis the work that turns it into something a team can act on.
Sizes Are Local to the Team
This constraint gets violated constantly, usually by well meaning managers.
A Medium reflects one team's context: their familiarity with the codebase, their skills, their definition of done, their environment. Another team looking at identical work may reasonably call it a Large.
That does not make either wrong. Relative estimation is calibrated against the team's own reference points, and those references differ.
The consequence is that aggregating sizes across teams produces meaningless numbers. Comparing how many Mediums each team completed is worse than meaningless, because it invites performance conclusions from data that cannot support them.
When leadership needs a cross team view, the honest answer is that estimates do not provide one. What does work is measuring delivered outcomes, cycle time or throughput of finished items, none of which depend on comparing subjective sizes between groups.
Protecting a team from having its estimates used this way is a genuine part of the Scrum Master role, and it is one of the more difficult conversations the job involves. It is covered practically in CSM Certification Training.
A Worked Session
A concrete run through, because the mechanics read as obvious and the facilitation is where sessions go wrong.
A team has fourteen epics to size ahead of quarterly planning. The Product Owner has written a short description for each. Nobody has analysed them in detail, which is the correct state for this exercise.
The Scrum Master opens by picking a reference. The team recently delivered a notifications feature that everyone remembers clearly, and they agree that is a Medium. It stays visible for the whole session.
The first epic is a search improvement. Sizes come back M, M, L, M, L. Close enough that no discussion is needed, and the team settles on M after ten seconds of conversation. Total time, under a minute.
The third epic is a payments integration. Sizes come back S, M, XL, XL, M. That spread is the reason the session exists. The two people who said XL work on the payments area and know about a compliance requirement the others had not considered. The person who said S was thinking only of the front end change. Five minutes of discussion, re-size, and it lands at XL.
Two things happened there. The item got sized, which was the stated purpose. More importantly, a compliance requirement surfaced during a planning conversation rather than mid Sprint three months later. That second outcome is worth considerably more than the label.
Because it came back XL, the Product Owner marks it for breakdown before it goes anywhere near a Sprint.
The remaining eleven epics take about thirty minutes, with two more meaningful discussions and the rest settling quickly. Total session, under an hour, and the team leaves with a sized backlog and three assumptions corrected.
The pattern worth noticing is that most items size quickly and a small number generate real discussion. A session where everything is unanimous usually means people are not thinking, and a session where everything is contested usually means the items are too vague to size at all.
Facilitating Disagreement Well
The outliers are the point, and how they are handled determines whether people keep contributing honestly.
Ask the extremes to explain, in either order. A common convention is highest and lowest, and it matters that both speak. If only the high estimate is questioned, people learn that large sizes require justification and small ones do not, and estimates drift downward over time.
Ask what they know, not why they are wrong. The person who said XL is not making an error, they are holding information. Framing the question as what are you seeing that others might not gets that information out. Framing it as why do you think it is so big produces defensiveness.
Watch for the quiet resize. After a discussion, people sometimes change their estimate to match the group rather than because they were persuaded. If someone shifts from XL to M without saying what changed their mind, that is worth a gentle check. Unspoken disagreement resurfaces later as a missed estimate.
Do not average. Averaging M and XL to produce L discards the disagreement rather than resolving it. The gap existed for a reason and averaging buries it.
Accept unresolved cases. Occasionally a team genuinely cannot agree because the item is not understood. The right output is not a size but a decision to investigate, timeboxed, and to size it afterwards.
The wider skill here is running a conversation where people say what they actually think, which is the same skill that makes retrospectives and refinement work. It is the least technical part of estimation and the part that most determines its value. Developing it properly is a substantial component of CSM Certification Training.
Common Mistakes
Beyond the conversion trap, six patterns cause most of the trouble.
No reference item. Without an agreed Medium, sizes drift as the session progresses and early items are sized against nothing.
Sizing by duration. Someone asks how many days a Medium is, and the team starts estimating time rather than relative size. Once that happens the technique has become time estimation with unusual labels.
Sizing items nobody understands. If the team cannot describe what an item involves, the size is a guess. The correct response is to timebox an investigation, not to produce a number.
The most senior voice deciding. Simultaneous reveal exists specifically to prevent this. Where a technical lead's opinion consistently becomes the size, the session is theatre.
Sizes used as commitments. An estimate is a forecast under uncertainty. Once sizes are treated as promises, they inflate, because inflating is the only defence available to the team.
Never recalibrating. Team capability changes. Work that was a Large eighteen months ago may be a Small now, and a team that never revisits its references gradually loses the meaning of its own scale.
Sizing alone and presenting the result. A single person sizing the backlog before the session removes the comparison conversation entirely, and that conversation is where most of the value sits. The sizes may be perfectly reasonable and the team still learns nothing, because nobody had to reconcile a differing view. It also quietly transfers ownership of the estimate away from the people who will do the work, which tends to show up later as reduced commitment to the plan.
How T-Shirt Sizing Fits the Wider Backlog
The technique is most valuable as part of a progression rather than as a standalone practice.
Far out work, at the epic or initiative level, gets t-shirt sized. That is enough to decide sequencing, spot the items that are far too large, and have a sensible roadmap conversation without pretending to precision.
As items approach being worked on, they get broken down and refined. Understanding improves, and at that point finer estimation becomes possible and useful. Many teams switch to story points here, often using planning poker to run the conversation.
By the time work enters aSprint, items should be small and well understood enough that estimation is quick and largely uncontroversial.
The estimation technique should match the level of understanding available. Applying fine grained numeric estimation to something nobody has analysed produces a number that looks rigorous and is not. Applying coarse sizes to a well understood story wastes the understanding the team has.
Getting that match right is more useful than any argument about which technique is superior.
Should You Estimate At All?
Worth addressing, since some teams have moved away from estimation entirely.
The no estimates position argues that estimation consumes time, produces unreliable numbers, and that simply slicing work into small consistent pieces and counting throughput forecasts better. There is real evidence behind this, and for teams with genuinely uniform work it holds up.
The counter argument is that estimation conversations surface disagreement. When one person says Small and another says XL, the discussion that follows frequently uncovers a misunderstanding that would have caused problems later. The number is disposable, the conversation is not.
T-shirt sizing sits usefully between the positions. It is cheap enough that the cost objection largely disappears, while still producing the comparison conversation. A team spending an hour sizing a quarter's worth of epics is not investing heavily in estimation, and it gets the discussion.
For most teams the practical answer is to estimate lightly at the level where decisions are being made, and to stop estimating anything that does not inform a decision. Estimates produced for reporting purposes, which nobody uses to decide anything, are the ones worth removing first.
Using T-Shirt Sizing With Stakeholders
One of the strongest arguments for the technique has nothing to do with the team, and it is rarely mentioned.
Business stakeholders engage with sizes in a way they do not engage with story points. Told an initiative is 21 points, a commercial director has no basis for reacting. Told it is an XL and the two things beside it are Smalls, the trade off becomes immediately legible.
That legibility changes prioritisation conversations. Sequencing arguments improve considerably when everyone in the room can see that one request is several times the size of the alternatives, and it removes a common failure where stakeholders push for everything because nothing appears to cost anything.
Two cautions apply.
Sizes are not effort in days, and someone will eventually ask how long an XL takes. The honest answer is that it depends on what else the team is doing, and that the size indicates relative scale rather than a duration. Giving a day figure to be helpful converts the size into a commitment, and it will be quoted back later.
Stakeholders should not set the sizes. They can and should explain what an item involves and what value it carries, but the size is a judgement about effort, and that belongs to the people doing the work. Where business stakeholders size the work, items they favour become mysteriously small.
Used carefully, sizing sessions with stakeholders present are among the more productive conversations available, because the two groups are finally discussing the same trade off with the same information. It pairs naturally with structuredprioritisation techniques, where relative size is one input into the sequencing decision.
Frequently Asked Questions
1. What is t-shirt sizing in agile?
A relative estimation technique using labels such as XS, S, M, L and XL rather than numbers. The team compares items against each other rather than estimating how long each will take.
2. How does t-shirt sizing differ from story points?
T-shirt sizing is coarser, faster and uses labels that do not add up. Story points are numeric, support velocity and forecasting, and suit smaller well understood items. They generally operate at different levels of detail.
3. Can you convert t-shirt sizes to story points?
You can, and it usually removes the point of using sizes. Mapping S to 2 and M to 5 recreates story points with extra steps and asserts a precision the underlying judgement does not support.
4. When should a team use t-shirt sizing?
For epics, features, roadmap planning and initial backlog triage, where items are large and not yet well understood. It is less suitable for sprint level stories or for capacity forecasting.
5. Can you calculate velocity from t-shirt sizes?
Not reliably. Sizes are labels rather than numbers and do not aggregate meaningfully. Teams needing velocity should use story points for sprint level work.
6. What does XL mean in t-shirt sizing?
In practice it is a signal rather than an estimate. An XL item is too large to plan and needs breaking into smaller pieces that each deliver value.
7. Can sizes be compared between teams?
No. Sizes are calibrated to one team's context, skills and reference points. Comparing them across teams produces conclusions the data cannot support.
8. How long should a sizing session take?
Twenty to thirty items in an hour is a reasonable pace. Sessions that run much longer usually mean the items are not understood well enough to size, which points at refinement rather than estimation.
9. What if the team has no reference item?
Pick recently completed work everyone remembers and agree it is a Medium. A shared reference is what makes the estimates relative rather than arbitrary, and without one the early items get sized against nothing.
10. Should a team recalibrate its sizes over time?
Yes. Capability changes, and work that was a Large two years ago may be a Small now. Teams that never revisit their reference points gradually lose the meaning of their own scale.
11. Who should take part in a sizing session?
The people who will do the work. A Product Owner or business stakeholder can explain what an item involves, but the size is the team's judgement, since they are the ones who will build it.
Closing Thoughts
T-shirt sizing earns its place by being fast and honest about uncertainty. Where work is large and poorly understood, a rough comparison is genuinely more useful than a number that implies knowledge nobody has.
The technique fails in two predictable ways. Teams apply it to work that is already well understood, where finer estimation would serve better. Or they convert sizes into numbers to make them add up, which discards the reason for using sizes in the first place.
A useful habit for any team adopting it: agree the reference item, write it somewhere visible, and revisit it every few months. Most of the problems described above trace back to a scale that drifted because nobody maintained the anchor.
The wider principle is that the estimation technique should match how much is actually known. Rough understanding, rough estimate. Detailed understanding, finer estimate. No understanding at all, investigate rather than estimate.
If you want to build the facilitation skills that make sizing sessions productive rather than performative,CSMCertification Training from an accredited Scrum Alliance provider covers estimation, refinement and running these conversations well. Request the full course curriculum to see the agenda and upcoming dates, or check your current understanding with our free CSM practice test.


























