Autumn forest interior with strong sun shafts raking across a rust-coloured leaf-covered floor

Our method

How We Test Survival Gear

This page is the standard itself, not a description of one. It sets out how we test survival gear at rationalsurvivor.com: how an item enters the queue, how long it is used before we write a word, which conditions get recorded, how timed results are reported, and what we refuse to test at all.\n\nRational Survivor is a new publication. That means most of our tests run on a single sample, and a single sample can only support certain kinds of claim. Where that is the case, the article says so in the same sentence as the result rather than in a footnote.\n\nEverything below is checkable against what we publish. If a piece on this site breaks one of these rules, tell us — the last section explains exactly how to make that stick.

The Short Version

Seven commitments. Every gear piece on this site is held to all of them.

  • We use it before we write about it. Each category has a minimum use period, stated in the article. Below that threshold, the piece is labelled a first look and carries no verdict.
  • We publish conditions. Temperature, precipitation, wind, elevation, water source, operator state. A result without conditions is an anecdote.
  • We publish spreads, not best attempts. Median, slowest attempt, number of attempts, number of failures.
  • We publish failures, including ours. A failure that happened once is reported as having happened once, and it stays in the article afterwards.
  • We state sample size every time. When it is one, it says one.
  • Safety-critical numbers are not ours. Boil times, bleach ratios, food spoilage windows and carbon monoxide distances come from published agency guidance, linked to the source.
  • We separate what we measured from what we read. Our figures carry conditions. Agency figures carry links. They never share a sentence without attribution.

The reasoning behind these, and the wider rules on sourcing and corrections, sit in our editorial policy. Finished tests are collected under gear tested.

How Gear Enters Testing

An item enters the queue when it maps to a task we already teach, not when a press release lands.

Those tasks come from the drills in our guides: make one gallon (3.8 L) of doubtful water safe to drink in under 30 minutes, get a self-sustaining flame in rain, carry 72 hours of supplies 3 miles (5 km) on pavement, keep a medical device running through an eight-hour night. If a product does not address a task like one of those, it does not get tested — however well it sells.

Where the gear comes from.

  • We buy at retail wherever the budget allows, from ordinary consumer listings, so we get the unit you would get.
  • Manufacturer samples are accepted for testing and labelled as samples in the article, every time. Accepting one guarantees nothing: not coverage, not a publication date, not a rating.
  • No payment is accepted for a review, a ranking, a place in a buying guide, or the removal of a finding. Where a link earns commission, our affiliate disclosure explains what that does and does not change.
  • Unsolicited items sent without an email exchange first may not be tested at all.

Before the first use, we write down the manufacturer's claims verbatim — rated temperature range, flow rate, runtime, capacity, duty cycle, service interval. The test is then run against that claim. This matters because a product that never claimed to work at 10 °F (-12 °C) should not be criticised for failing there, and a product that did claim it should be held to it.

Minimum Use Before We Publish

Most gear works on day one in a kitchen. The useful information arrives later. These are the floors we hold ourselves to before an item can carry a verdict or appear in a recommendation.

  • Packs and carry systems: 60 days of real use and at least 50 miles (80 km) loaded, including three walks at full 72-hour kit weight — roughly 10 percent of body weight, about 15 lb (7 kg) for a 155 lb (70 kg) adult.
  • Cutting tools: 30 days of use and one full dulling-and-resharpening cycle, so we can report how the edge behaves as it degrades rather than how it arrived.
  • Fire tools: a minimum of 30 ignition attempts across at least three tinder types, of which at least five are in falling rain or on saturated ground. See starting a fire in wet conditions.
  • Water treatment hardware: at least 30 field fills from a real source, including one session at or below 40 °F (4 °C), with flow rate tracked across the period.
  • Lights and batteries: ten full charge-and-discharge cycles, one of them at or below freezing, 32 °F (0 °C).
  • Stoves: 30 burns, including cold-canister starts and a boil in wind.
  • Shelter and sleep systems: 15 nights outdoors, at least one in sustained rain and one at the bottom of the rated temperature range.
  • First aid kits: unpacked, repacked and inventoried three times, including one access test one-handed and one in gloves and darkness, plus a full expiry audit. Contents are judged against the packing list in our wilderness first aid kit guide.

If we have not met the floor, the article says so at the top and is published as a first look: observations, specifications, claims recorded, no verdict and no place in any recommendation.

The Conditions We Write Down

A time or a result means nothing without the conditions attached to it. The following are recorded at the test site, at the time, on paper or in a phone note before the session ends — never reconstructed from memory afterwards.

  • Air temperature at the start and end of the session, read at the site rather than taken from a weather app.
  • Precipitation: none, drizzle, steady rain, wet snow or dry snow, plus how long it had been falling and whether surfaces were already soaked when we started.
  • Wind: sustained and gusting. Where we do not have a handheld reading, we use the nearest National Weather Service observation and name the station and its distance from us, so you can judge how much to trust it.
  • Elevation to the nearest 100 ft (30 m). It changes boiling point, stove output and how hard the work feels.
  • Water source and state: tap, rain barrel, pond, stream or meltwater; clear or turbid; temperature.
  • Operator state: bare or gloved hands, wet or dry, daylight or headlamp, first attempt of the day or the tenth, and whether the operator had already walked several miles.
  • Consumable state: battery charge, fuel level, canister temperature, cartridge hours used.

The corollary matters as much as the list. A condition we did not record is never implied later. If we did not measure wind, the article says wind was not measured — it does not say "in calm conditions".

How Timed Tasks Are Run — and Reported as a Spread

Timed tasks are where gear writing usually goes wrong, because one lucky attempt makes a better headline than five honest ones.

How a task is run.

  • The start and stop points are defined in writing before the first attempt. For example: the clock starts when a hand touches the pack and stops when the flame has survived two minutes unattended. Ambiguity in the stop condition is what lets a number drift.
  • A minimum of five attempts per operator per condition.
  • The first attempt counts. No rehearsal rep is discarded, because your first attempt will not be discarded either.
  • The same operator runs a task across conditions wherever possible. Where more than one person takes part, results are separated by operator rather than pooled.

How it is reported.

  • Median, slowest attempt, number of attempts, and number that failed outright.
  • A first-attempt figure reported separately from the practised figures, because a person who has run a drill 40 times is not you.
  • No best times. A best attempt is a story about one repetition, and it is the number most likely to get someone into trouble when they plan around it.

We do not re-run a task until a number looks good and then print that one. If a session produces an ugly spread, the ugly spread is the result.

What Counts as a Failure

Failure is defined before testing starts, in four classes, so that the definition cannot move to suit the outcome.

  • Hard failure. The item stops doing its job and cannot be recovered in the field: a cracked housing, a dead cell, a snapped strap, a pump that will not prime. Published with the conditions and the elapsed use at the moment it happened.
  • Degraded performance. It still works, but worse: flow rate falls, runtime shortens in the cold, a zip catches, a valve weeps. Published with the trend over the use period.
  • User-defeatable failure. It failed because of how it was used, and a different technique fixes it. This gets published as a technique note rather than held against the product, because the same mistake is waiting for you.
  • Out of scope. It failed doing something it was never sold to do. Recorded for context, not counted against the item.

Three rules govern what happens next. A failure is written down at the time it occurs, not reconstructed later. A failure is never deleted because a replacement unit worked — if a manufacturer sends a replacement, the original failure stays in the article and the replacement is reported as a second sample. And a failure is never traded away: we do not remove findings in exchange for units, access or advertising.

Sample Size, Stated Plainly

Most tests on this site run on one unit. That is the honest position for a new publication, and pretending otherwise would undermine everything else on this page. So the sample size appears in the article, next to the result.

What a sample of one can legitimately show:

  • Whether the design works in the hand, in gloves, in the dark, when you are cold and hurried.
  • Whether measured behaviour matches the manufacturer's stated claim.
  • Failure modes that appear during ordinary use — which are worth knowing about even at n=1.
  • Whether the thing is the right shape for the job at all.

What a sample of one cannot show:

  • Reliability rates, manufacturing consistency or expected lifespan.
  • Whether a failure we saw is typical or a one-off.
  • How a product compares with a rival on durability.

So one broken sample is written as "this unit failed under these conditions after this much use", never as "this model is unreliable". The two sentences look similar and mean entirely different things.

For the same reason, we do not publish scores out of ten. A single number conceals exactly what the conditions section exists to reveal. A recommendation here is always scoped: for this task, at this temperature, for a person carrying this load. Where a question genuinely needs a fleet of units and several years to answer, the honest answer is that we do not know yet, and that is what we will write.

Agency Figures Versus Our Own Measurements

There are two kinds of number on this site and they are never blended.

Agency figures carry a link to the primary source. Our measurements carry conditions, a spread and a sample size. If a sentence contains both, it attributes both.

Safety-critical figures are never ours. We do not derive them from our own testing, and we do not round, simplify or "update" them. These come from published guidance:

When published guidance changes, the guidance wins over our older article. The page is corrected in place and carries a dated note saying what changed and whether the advice changed with it. Manufacturer claims are a third category again, always labelled as claims — including certification marks, which we report by the standard's name rather than restating as our own finding.

What We Do Not Test, and Why

A testing standard is defined as much by its limits as by its protocol. We do not run the following, and we will not imply that we have.

  • Pathogen removal. We have no accredited laboratory. We cannot verify that a filter removes bacteria, protozoa or viruses, so we never write that it does. We report the certification a product holds, by standard, and the manufacturer's claim, and we point you at EPA guidance for methods with published effectiveness. What we can test is flow rate, clogging, cold-weather behaviour, cleaning and field repair.
  • Medical procedures on people. First aid testing covers kits, packing, one-handed access, glove use, cold hands and expiry auditing. Techniques come from published guidance and hands-on training, not from us experimenting on volunteers. See the disclaimer.
  • Weapons and defensive tools. Outside our scope, and a long way from where most household risk actually sits.
  • Destruction for spectacle. Batoning a knife until it snaps produces video, not information about the use you will put it to.
  • Claims we cannot outlive. A publication this young cannot verify a 25-year shelf life. We report the manufacturer's figure, note how strongly storage temperature affects it, and defer to NCHFP on preservation.
  • Conditions we cannot exit safely. No untrained cold-water immersion, no deliberate carbon monoxide exposure, no burning fuel indoors to see what happens.
  • Licensed work. Generator transfer switches, gas appliances and fixed wiring are the domain of licensed trades. We test the portable equipment and the household plan around it, and say where the line is.

If you are deciding where to start rather than what to buy, start here is the better page — the editorial line on this site is that skills beat stockpiles, and most of our drills need no new gear at all.

How to Challenge a Result

A published standard is worth little without a route to enforce it. If a result here is wrong, we want to know, and corrections take priority over everything else in the inbox.

Write to hello@rationalsurvivor.com with Correction or Challenge at the front of the subject line, and include:

  • The page URL and the exact sentence or figure, quoted.
  • What you did and what happened, with your conditions: temperature, precipitation, wind, elevation, water source, operator, number of attempts.
  • The exact model, and any lot or date code, since a revision may explain the difference.
  • The source that contradicts us, where one exists — a superseded guideline, a manufacturer specification, a published figure.

What happens then:

  • If the disputed number is an agency figure, we re-check it against the primary source. If the guidance has moved, the guidance wins and the page is corrected.
  • If the number is ours, we re-run the task where the item and the conditions are still available to us, and publish the new spread alongside the old one rather than replacing it silently.
  • The page carries a dated note describing what changed. If a reader following the old version would have stored too little water, dosed a treatment wrongly or trusted a piece of gear further than it deserved, the note says exactly that instead of calling it a typo.
  • You get credited by name or initials if you want to be. Say so in your email.

What we will not do: remove an unfavourable finding because a brand asked, swap a spread for a flattering best attempt, quietly delete a page instead of correcting it, or treat "our lab says otherwise" as evidence without the conditions and sample size behind it.

Full routes and response times are on the contact page, and the standards this testing protocol sits inside are in our editorial policy.