Direct answer: name who the advice is for, the action it recommends, what it is compared with, the outcome that should change, the time window and the evidence that would count against it. If you cannot fill those fields, the advice may still be inspiring—but it is not ready to be treated as a tested claim.
Try the one-sentence template:
For [population], doing [action] rather than [comparison] will change [measurable outcome] by [time]. I would weaken or reject this claim if [falsifier] occurred.
The template does not make a claim true. It makes the missing evidence visible.
Why advice resists testing
Popular advice is often designed to travel across situations:
- Start with why.
- Let them.
- Just take action.
- Buy a boring business.
- Make an offer people cannot refuse.
- Never use a credit card.
That flexibility helps a phrase spread. It also lets the meaning change after the outcome is known. A success can be credited to the method; a failure can be blamed on using it incorrectly, not wanting it enough or applying it in the wrong context.
A testable claim creates a fairer rule. Before seeing the result, it specifies what should happen, for whom, under which conditions and what evidence would reduce confidence.
The BYBI Claim Test
Use the six core fields in order.
1. Population: who is the advice for?
“Everyone” is rarely an adequate answer.
Specify the people, businesses or situations to which the claim applies. Include material eligibility conditions where they affect the result.
Examples:
- First-time US founders starting employer businesses.
- Adults trying to initiate a clearly defined low-risk task.
- Managers communicating an established strategy to existing staff.
- Couples in a non-abusive relationship discussing a recurring conflict.
Population is where many transferability problems hide. Evidence from students may not apply to executives. Evidence from an expert-run intervention may not transfer to unsupported self-use. A method that helps with a low-stakes choice may be unsafe in an emergency.
2. Action: what exactly does the advice ask someone to do?
Replace a slogan with observable behaviour.
Weak:
Use the 5 Second Rule.
Stronger:
When a planned task cue occurs, count backward from five and begin the first physical action before starting another activity.
The stronger version can be taught, recorded and compared. It also reveals implementation details the slogan concealed.
3. Comparison: compared with what?
Improvement requires a reference point.
Possible comparisons include:
- Usual behaviour.
- No intervention.
- A different method.
- The same person’s baseline.
- A matched group.
- A different sequence or message.
Without a comparison, change may reflect time, regression to the mean, selection, concurrent actions or the fact that motivated people chose to use the method.
Not every personal decision needs a randomised trial. Every causal claim still benefits from asking what else could have produced the outcome.
4. Outcome: what should change?
“Better,” “successful,” “confident” and “inspired” need an operational meaning.
Choose an outcome that can be observed and that matters to the decision:
- Task initiated within two minutes.
- Task completed by the agreed date.
- Voluntary staff retention after twelve months.
- Qualified enquiries per week.
- Contribution margin after refunds and variable fulfilment costs.
- Conflict frequency recorded over eight weeks.
Do not silently replace the important outcome with an easier surrogate. Starting a task is not the same as completing it. Revenue is not profit. Feeling persuaded is not evidence that the proposal works.
5. Time: by when?
A claim without a time window can remain “not yet disproved” indefinitely.
Specify:
- When the action begins.
- When the outcome is measured.
- Whether persistence or follow-up matters.
- Which update would trigger a new review.
Short-term and long-term effects can differ. A method may create immediate relief and increase avoidance later. A promotion may raise this week’s cash and damage retention. The time window is part of the claim, not a footnote.
6. Falsifier: what would reduce your confidence?
A useful falsifier is evidence you would accept even if you wanted the claim to be true.
Examples:
- No meaningful difference from the comparison across adequately powered direct studies.
- Benefits disappear outside an expert-run setting.
- The result depends on excluding refunds, costs or failed users.
- The effect reverses for an important subgroup.
- Independent records do not reconcile a first-party number.
- Material harms outweigh the benefit for the target population.
The falsifier does not need to prove the opposite. It defines what would make you lower confidence, narrow the population or change the decision rule.
Add two credibility fields
The six fields make the claim testable. Two additional fields help you decide how much evidence it needs.
Source trail
Identify:
- The best primary source for the exact wording.
- The best independent evidence for the outcome.
- A credible countercase or limitation.
- The date the evidence was checked.
First-party sources are often strongest for what someone said, sold or did. They are rarely sufficient for typical outcomes or general transferability.
Incentive and cost of being wrong
Record the source’s material incentives and the reader’s downside.
A reversible, low-cost personal prompt can justify a small self-test. A medical, legal, financial or high-cost business decision requires much stronger evidence and qualified advice. Evidence standards should rise with the consequence of error.
Worked example: “Start with why to inspire action”
Raw advice
Start with why to inspire action.
Testable version
For employees receiving a presentation about an already approved organisational initiative, opening with a short authentic purpose statement rather than an otherwise identical task-first opening will increase the proportion who correctly recall the initiative and voluntarily complete the first assigned step within seven days. Confidence would weaken if the effect did not replicate, if it appeared only when the presenter was also the framework’s advocate, or if the purpose statement was not practised by leadership.
The testable version exposes choices:
- Which employees?
- Which presentation?
- What does “inspire” mean?
- What is the comparison?
- What happens within seven days?
- What result would count against the claim?
It also leaves room for a different verdict in an emergency, regulated disclosure or new product test where safety, evidence or customer need should come first.
Worked example: “Most startups fail”
Raw claim
Most startups fail.
Testable version
Among US private-sector employer establishments born in a named year, fewer than half will remain operating five years later, using the Bureau of Labor Statistics survival definition.
Now the claim can be checked against an identified cohort and outcome. It cannot be silently applied to every venture-backed company, side project or solo founder.
Evidence quality comes after claim clarity
Making a claim testable is the beginning of evaluation, not the end.
Evidence confidence can fall because of:
- Risk of bias.
- Inconsistent findings.
- Indirect populations, actions or outcomes.
- Imprecise estimates.
- Publication or reporting bias.
The Cochrane GRADE framework formalises these considerations for bodies of intervention evidence. BYBI adapts the underlying discipline—especially outcome specificity, directness and uncertainty—without pretending that every consumer claim is a healthcare review.
The practical sequence is:
- Make the claim specific.
- Find the best direct evidence.
- Inspect bias, consistency, directness, precision and missing evidence.
- Separate evidence quality from fit for this reader.
- State what would change the verdict.
Match the test to the type of claim
Not every claim asks the same evidence question.
| Claim type | Example | What a useful test needs |
|---|---|---|
| Descriptive | “3,000 companies contact us each month” | Defined count, dates, deduplication rule and documentary source |
| Comparative | “This method works better than the alternative” | Same population, outcome and period for both options |
| Causal | “The countdown caused more action” | A credible comparison that reduces alternative explanations |
| Predictive | “Most firms like this will close” | Defined cohort, outcome horizon, calibration and out-of-sample performance |
| Transferability | “This business model works in any industry” | Boundary conditions, varied settings, costs, capability and failure cases |
| Value or preference | “This is the best framework” | Criteria and audience values; evidence cannot decide the values by itself |
Labelling the type prevents a common error: using descriptive evidence—such as one company’s history—to support a broad causal or transferable conclusion.
Run a safer personal test
When the advice is low-risk and reversible, a structured personal test can be useful even when population evidence is limited.
Before the test
- Choose one behaviour and one outcome.
- Record a baseline under ordinary conditions.
- Decide how long the test will run.
- Write the minimum change that would matter.
- List harms, costs or stop conditions.
- Decide what you will do if the result is mixed.
During the test
Record the action and outcome consistently. Do not count only the days you remembered the method. Note concurrent changes—new deadlines, medication, staff, advertising or life events—that could explain the result.
After the test
Compare the result with the baseline or alternative. Ask whether the change was large enough to matter, whether it persisted and whether the method created new costs. Treat one personal result as evidence for your decision, not a universal claim about everyone else.
When not to self-test
Do not use a casual personal experiment to replace urgent, qualified medical, legal, financial or safety advice. Do not test a method by exposing another person to undisclosed risk or withholding a duty of care. High-stakes claims need an evidence and review path proportionate to the harm.
Turn one slogan into a claim in under five minutes
Use this quick version when you encounter a video, article or sales page:
- Copy the exact sentence.
- Circle the verb: what is supposed to happen?
- Add the missing “for whom?”
- Add “compared with what?”
- Replace the vague result with one observable outcome.
- Add a date.
- Write one result that would lower confidence.
- Record who benefits if you believe it.
If the sentence becomes awkward, that is useful. Awkwardness often reveals that one portable slogan contains several different claims. Split them and test the decision-relevant one first.
Copy-and-use Claim Test worksheet
| Field | Your answer |
|---|---|
| Raw advice or claim | |
| Population | |
| Exact action or exposure | |
| Comparison | |
| Measurable outcome | |
| Time window | |
| Falsifier or confidence-lowering evidence | |
| Best primary source | |
| Best independent source | |
| Incentive or conflict | |
| Cost of being wrong | |
| Safer decision rule |
Final sentence
For [population], [action] compared with [comparison] is expected to change [outcome] by [time]. The best current evidence is [source and confidence]. I would narrow or reject the claim if [falsifier]. Because the cost of being wrong is [low/medium/high], the safe action is [decision rule].
Next step: complete the worksheet for the advice that is closest to changing your behaviour or spending. Save the source trail and falsifier before looking for more supporting examples. Subscribe only if you want the Claim Test and evidence ledger when the tool changes.
Common mistakes
Making the claim so narrow that it no longer matters
Specificity should protect the decision, not hide behind a trivial outcome. “People can count backward from five” is testable and irrelevant to the claim that the rule changes lives.
Choosing only a success metric
Include costs, harms, persistence and the people who do not benefit where they matter.
Treating no evidence as evidence of no effect
If direct evidence is absent, the honest conclusion is usually “uncertain” or “not directly tested,” not “proven false.”
Changing the outcome after the result
Decide in advance whether initiation, completion, profit, wellbeing or another outcome carries the claim.
Ignoring fit
A well-supported average effect can still be a poor fit for a person with different constraints, values or risks.
BYBI verdict
A claim becomes decision-useful when it is specific enough to be wrong. Population, action, comparison, outcome, time and falsifier are the minimum useful structure. The Claim Test improves transparency; it does not replace domain expertise, direct evidence or judgement about consequences.
Confidence: high for the principle of testability; moderate that this exact interface improves reader decisions until usability testing and error analysis are complete.
What would change this verdict?
The tool should be revised if reader testing shows that people consistently confuse “testable” with “true,” omit harms and costs, or use the worksheet to create false precision. It should become stronger if repeated testing shows improved claim specificity, source selection and confidence calibration.
Sources and update record
- National Academies/NCBI Bookshelf, Scientific Principles and Research Practices.
- University of California, Berkeley, Understanding Science.
- Cochrane, Completing Summary of Findings tables and grading the certainty of the evidence.
- Cochrane, Interpreting results and drawing conclusions.
- US National Institute of Standards and Technology, Measurement Uncertainty.
Last evidence check: August 22, 2026. Review after user testing, material changes to the BYBI evidence policy or relevant updates to cited standards.
Frequently asked questions
Is a testable claim automatically true?
No. Testability makes it possible to gather evidence that increases or reduces confidence.
Does every piece of advice need a controlled experiment?
No. The evidence burden should reflect the type of claim and cost of being wrong. A low-cost reversible prompt can justify a personal test; high-stakes causal advice requires stronger direct evidence and qualified review.
What is the difference between a claim and an opinion?
A claim asserts something that can be supported or challenged by evidence. An opinion may express a preference or value. Opinions can contain factual claims that still need testing.
What if there is no direct evidence?
Label the evidence as indirect or absent, lower confidence and use a safer, narrower decision rule. Do not convert adjacent research into direct validation.
What makes a good falsifier?
It is evidence you would accept as a reason to lower confidence, narrow the claim or change the decision—even if you prefer the claim to be true.

