What Forensic Evidence Actually Proves (And What It Doesn't)
Published August 14, 2026 Β· 6 min read
Television has taught two generations that forensic evidence is a machine that turns physical traces into certainty. A fingerprint means this person. A hair means this person. A bite mark means this person. The technician runs it, the screen flashes, the case closes.
Actual forensic science is enormously more varied in quality than that, and the differences between disciplines are not small. Some are grounded in measurable statistics and hold up under scrutiny. Others were accepted by courts for decades on nothing more than practitioner confidence, and have since been substantially discredited.
Two major reviews established this publicly: the US National Academy of Sciences report Strengthening Forensic Science in the United States (2009), and the President's Council of Advisors on Science and Technology (PCAST) report on feature-comparison methods (2016). Both reached an uncomfortable conclusion β that with the notable exception of nuclear DNA analysis, many forensic disciplines had not been validated to the standard their courtroom testimony implied.
Here's the practical version, discipline by discipline, and what it means for reading any mystery.
The core distinction: identification vs comparison
Almost every argument about forensic reliability comes down to one distinction.
Statistical identification compares a trace against a population and produces a number: the probability that a randomly selected unrelated person would match. DNA does this. The number can be astronomically small, and β crucially β it's derived from measured population data rather than from the examiner's judgement.
Feature comparison puts two things side by side and asks a trained examiner whether they're similar enough to have come from the same source. Fingerprints, toolmarks, handwriting, tyre treads, hair and bite marks all work this way. The examiner's conclusion may well be correct. What it usually lacks is a validated error rate, a defined threshold for "enough," and protection against knowing what answer the investigation wants.
That second category is where the problems live β not because the methods are worthless, but because the language used to report them ("a match," "to the exclusion of all others") claimed far more than the underlying science supported.
Discipline by discipline
Nuclear DNA β the strongest
DNA profiling is the one discipline both major reviews treated as scientifically sound. It has a defined statistical foundation, published population frequencies, and known limitations.
Its real problems are not about the chemistry. They're about interpretation and provenance:
- Mixtures. A sample containing DNA from three or more people is far harder to interpret than a single-source sample, and analysts examining the same mixture have been shown to reach different conclusions. Probabilistic genotyping software has improved this, and has introduced its own arguments about validation and transparency.
- Transfer. Modern methods can profile tiny amounts of material, which means DNA can be detected from indirect contact β you shake someone's hand, they touch a surface, your DNA is on it. Sensitivity increased faster than the interpretive framework for what a trace means.
- Presence isn't participation. DNA on an object establishes contact, not when, how, or in what circumstances. A profile on a weapon is compatible with an attack and with having handled it a week earlier in the kitchen.
The useful mental model: DNA reliably answers whose. It does not answer when, how, or why.
Fingerprints β good, but not infallible
Latent print comparison is a genuine skill with a real basis, and error rates are low. But "infallible" β a claim made in courts for most of the twentieth century β was never supported.
The Brandon Mayfield case made this concrete. After the 2004 Madrid train bombings, an Oregon lawyer was linked to a print from the scene by multiple FBI examiners and detained, before Spanish authorities identified a different man. The FBI's own review afterwards pointed to circular reasoning and the influence of contextual information on the examiners.
Two structural weaknesses persist. Crime-scene latents are partial, smudged and distorted, which is a fundamentally harder task than comparing two clean rolled prints. And the process has historically been vulnerable to contextual bias β examiners who know the suspect confessed, or that other evidence points the same way, are more likely to declare a match. Blind verification procedures were introduced specifically to counter this.
Firearms and toolmarks β contested
The claim is that a barrel or a tool leaves individual, reproducible marks that a trained examiner can trace back to one specific weapon. PCAST concluded the foundational validity of this was supported by only limited evidence, and courts have since restricted how confidently examiners may phrase their conclusions.
The direction of travel is away from "this bullet was fired from this gun to the exclusion of all others" and toward statements about degrees of consistency.
Hair microscopy β largely superseded
Microscopic hair comparison was used for decades to link suspects to scenes. In 2015 the FBI, working with the Innocence Project and the National Association of Criminal Defense Lawyers, published the results of an internal review of historical cases and found that examiners had given scientifically invalid testimony in the overwhelming majority of the trials reviewed β frequently overstating what a microscopic similarity could establish.
Mitochondrial DNA testing of hair is a real and useful technique. The microscopic comparison that preceded it is not a means of individual identification.
Bite mark analysis β discredited
Forensic odontology's claim that human dentition is unique and that skin reliably records it has not survived examination. PCAST found the method scientifically invalid, studies have shown poor agreement between examiners on whether an injury is even a bite mark at all, and a series of convictions built on bite mark testimony have been overturned.
It is the clearest case of a discipline achieving courtroom acceptance long before anyone tested whether it worked.
Blood spatter analysis β depends entirely on the claim
Basic bloodstain interpretation rests on real fluid dynamics: droplet shape indicates angle of impact, and convergence can indicate an origin area. That much is defensible.
The problems arise with elaborate reconstructions β sequences of events, positions of participants, the specific weapon β presented with a confidence the underlying physics doesn't support. The NAS report criticised the field's lack of standardisation and the variability in analysts' training.
Arson investigation β rebuilt from the ground up
For decades, fire investigators identified arson using indicators β "crazed glass," "pour patterns," "alligator charring" β that were later shown by fire-science research to occur in accidental fires too, particularly after flashover. This was not a marginal correction; it invalidated a substantial body of past casework, and at least one execution in the United States has been the subject of sustained expert dispute on exactly these grounds.
Modern fire investigation is governed by a scientific-method standard (NFPA 921) which explicitly rejects the older indicator-based approach.
The two biases that do the most damage
Contextual bias. Examiners are usually told things they don't need: that the suspect confessed, that a witness identified them, that other tests agree. Research on forensic decision-making shows this shifts conclusions in ambiguous comparisons. The countermeasure β linear sequential unmasking, where an examiner documents their analysis of the crime-scene trace before seeing the reference sample β is straightforward, and adoption has been uneven.
Confirmation through convergence. When five weak pieces of evidence all point one way, a case feels overwhelming. But if all five were produced by examiners who knew the theory of the case, they aren't independent β they're five expressions of the same assumption. A case can look corroborated when it's actually just repeated.
This is the more useful of the two for any detective, real or fictional: evidence only corroborates if it was generated independently.
The CSI effect
The popular claim is that crime dramas raise jurors' expectations so they acquit without forensic evidence. Empirical research on this has been mixed, and it's less well-evidenced than the confident way it gets discussed.
There's a second, better-supported version worth knowing about, sometimes called the "tech effect" or the reverse CSI effect: television inflates confidence in what forensic science can do. Jurors, and readers, arrive believing that a match is a certainty rather than a probability, and that every trace can be recovered, tested and resolved. That prior does more damage than scepticism does.
How to read forensic evidence properly
Four questions handle nearly every claim, in fiction and in the real thing:
1. Is this identification or comparison? If a number is attached and it came from population data, you're in the strong category. If the conclusion is an examiner's judgement of similarity, ask what the error rate is.
2. What does presence actually establish? A trace on an object proves contact between that person and that object at some unspecified time. Getting from there to "committed the crime" requires additional argument, and that argument is where cases are won and lost.
3. Was the examiner blind? Did they know what the investigation wanted the answer to be? If yes, treat the result as one opinion, not one measurement.
4. Are the corroborating pieces independent? Five results produced by people who all knew the theory of the case are one result.
For mystery readers and players, that fourth question is the most productive. It's the structural weakness a well-built case exploits: a suspect who looks buried under converging evidence, all of which flows from a single upstream assumption that nobody re-examined. Break the assumption and the whole edifice goes. The Perfect Witness is built on precisely that shape β evidence that never lies, written by hand β and The Nullifier works the inverse, where the absence of evidence is itself the manufactured thing.
The same logic makes better fiction. A mystery whose solution turns on the interpretation of physical evidence, rather than on whether it exists, is fairer and far more satisfying than one where a lab result simply arrives and settles the matter. Our writing guide treats the reveal as a re-reading rather than a delivery of new facts, and forensic evidence works best under exactly that constraint.
Frequently asked questions
Is forensic science reliable?
It depends entirely on the discipline. Nuclear DNA analysis has a strong statistical foundation. Fingerprint comparison is a real skill with low but non-zero error rates. Bite mark analysis has been found scientifically invalid, and several other feature-comparison methods have been found to lack the validation their courtroom testimony implied. "Forensic science" is not one thing.
Can DNA evidence be wrong?
The chemistry is rarely wrong; the inference frequently is. Complex mixtures are genuinely difficult to interpret, transfer means DNA can arrive on an object without its owner ever being present, and a profile establishes contact rather than participation. Contamination and sample-handling errors are the other real-world failure mode.
Are fingerprints unique?
The working assumption that friction ridge detail is highly individual is well accepted. The practical question is different: whether a partial, distorted crime-scene latent contains enough usable detail to identify one person, and whether the examiner comparing it was insulated from knowing what answer was wanted. Both have produced documented errors.
What is the CSI effect?
The claim that crime dramas change juror expectations of forensic evidence. Research on whether it causes acquittals is mixed. The better-supported concern runs the other way β that popular depictions inflate belief in what forensic science can deliver, making a "match" sound like certainty when it's a probability.
How should this change the way I read mysteries?
Treat physical evidence as a constraint rather than an answer. It narrows the field; the deduction still has to be done. The best mysteries β and the best real investigations β turn on what a piece of evidence means rather than on its existence, which is also the difference between a satisfying reveal and a lab report. If you want to practise reading evidence that way, browse the cases and see how far physical findings actually get you before you have to start interrogating people.