The Rubric: What It Is, Why It Works, And Why It Cannot Be Sanitized
PART I: WHAT THIS IS
This is a three-test filter for distinguishing science from sophistry. It applies the same three tests at two layers: the study (did this specific experiment test its claim honestly) and the paradigm (does the framework the study lives inside allow itself to ever be wrong).
The tests are not opinions. They are not heuristic guidelines. They are logical conditions that must hold for a claim to be anchored to reality rather than to its own framework. A claim that fails them is not necessarily wrong. It is ungrounded. It might be correct by accident. But it is not established by the evidence presented, because the evidence presented is the model talking about itself.
The verdicts are mechanical. Three passes means scientific. Three fails means unscientific. A mix means mixed. Any test marked Cannot Tell and no test marked Fail means Cannot Tell. This includes Cannot Tell on Test 1 alone with Pass on Tests 2 and 3. A pass on the other two tests does not rescue a test you cannot see. Any test marked Cannot Tell alongside at least one Fail means score what you can see. The missing information does not rescue a paper that is clearly failing on the tests you can score. The scorer does not adjudicate. The scorer applies the tests and the verdict falls out. There is no override. There is no "but this was published in Nature." There is no "but the consensus agrees." The structure of the experiment either meets the conditions or it does not.
If no theoretical framework is present, the paradigm layer is Not Applicable and the study verdict alone determines the outcome.
The paradigm layer has one edge case. Some studies do not operate inside any theoretical framework. They count visible events in defined groups and compare the counts. A child wheezing is visible. A parent bringing a child to the doctor is visible. A billing code is a filing label stuck onto a visible complaint after the fact, not a theoretical commitment. When a study's methodology is limited to counting visible events and comparing the counts, the paradigm layer is scored as Not Applicable. There is no model to be independent from, no model to falsify, and no model to be circular with. The three tests do not apply because there is nothing to test. The study verdict alone determines the final verdict.
A study inherits a paradigm only if its methodology depends on that paradigm's theoretical claims to function. Dependence means the study uses paradigm-dependent instruments to produce its observations, paradigm-defined tests to sort its groups, or paradigm-specific interpretations to derive its conclusions. A study that uses a paradigm's infrastructure, hospital records, billing codes, diagnostic nomenclature, as raw data for counting visible events does not operate inside that paradigm. Using a paradigm's records is not the same as operating inside its theory. The distinction is the same one the rubric already makes for instruments: the tool produces a signal, the model supplies the meaning. A billing code is the signal. The child's complaint is the meaning. The complaint is visible without the model. The code is just a label. If the study's comparison would mean the same thing to someone who rejects the paradigm entirely, the study is paradigm-independent.
The exception is only for visible events. An office visit, a death, a fracture, a fever, and a hospital admission exist without a model. A questionnaire score that the model names a depression case does not. The answers exist. The case does not exist without the psychiatric model. If any key variable is a model construct measured by a model-dependent instrument, the study operates inside that paradigm. The three tests then apply to both layers.
PART II: WHAT THIS IS NOT
It is not anti-science. It is the most aggressively pro-science position available, because it demands that science actually be science rather than perform science. A claim that cannot survive these three tests is not science regardless of how many journals published it, how many experts endorsed it, or how sophisticated the instrumentation appears. If that disqualifies most of published research, then most of published research is not science. The rubric does not flinch from that conclusion. It produces it mechanically.
It is not a conspiracy theory. It does not claim that scientists are lying, coordinating, or acting in bad faith. It claims something worse: that the structure of modern research produces fraudulent outcomes without any individual fraud being necessary. Scientists running PCR on samples they believe contain a virus are not lying. They are using a tool that was built from a model, calibrated to that model, and interpreted through that model, and then reporting the model's output as evidence for the model. Nobody lied. Nobody coordinated. The circularity is in the instrument design, not in the scientist's intent. The fraud is structural. It does not require a conspirator. It requires only that nobody asks whether the instrument is independent of the conclusion.
It is not relativism. It does not say all models are equally valid or that truth is subjective. It says the opposite: that a claim anchored to reality through independent observation, genuine falsifiability, and non-circular reasoning is more likely to be true than a claim anchored to its own framework. It takes sides. It says: a visible tumor shrinking is real. A fluorescence curve interpreted as viral load is an interpretation. Those are not equal. The first is observation. The second is the model narrating its own output.
It is not about credentials, consensus, or authority. A farmer who counts dead chickens and finds that the ones given clean water survived while the ones given contaminated water died has produced a scientific observation if the setup was falsifiable and non-circular. A Nobel laureate running a PCR assay on a chemically processed pellet has produced a model-dependent interpretation. The rubric scores the structure, not the source. A well-designed study by an unknown author passes. A circular study by a famous team fails.
It is not about being right or wrong. A claim can be wrong and still pass all three tests. Pasteur may have misidentified what was killing his silkworms, but his observation, count the worms in the infected vs. clean groups, was independent, falsifiable, and non-circular. The study was scientific even if the conclusion was imperfect. A claim can also be right and still fail. The terrain model of disease might be correct, but if a study "proves" it using a tool that was built from the terrain model's own definitions, the proof is circular. The claim might be true. The evidence is still invalid. The rubric scores the evidence, not the truth. Truth is not always available. Whether the evidence is structurally sound is always available.
PART III: WHERE THE TESTS COME FROM
Test 1: Independently Verifiable
Root: the empiricist requirement that knowledge be grounded in observation, not interpretation.
This traces to the oldest distinction in epistemology: between things you observe and things you infer. Hume separated matters of fact from relations of ideas. Matters of fact require sensory evidence. Relations of ideas are true by definition within a logical system. "2 + 2 = 4" is a relation of ideas. It is true within arithmetic. It does not require observation. "This solution contains a virus" is a matter of fact. It requires evidence. If the evidence is "the PCR machine produced a fluorescence curve and the model says that curve means virus," then the claim is not grounded in observation. It is grounded in interpretation. The observation is the curve. The claim is the virus. The model bridges the gap. Without the model, the curve is just a line on a screen and the virus is an assertion.
The test asks: can the observation exist without the model? A fever breaking is visible to anyone with eyes. A fluorescence curve being interpreted as "viral RNA detected" is visible only to someone who has already accepted the model that connects fluorescence curves to viral RNA. The first is an observation. The second is the model narrating its own output and calling the narration discovery.
The instrument interpretation rule follows from this. An instrument produces a signal. That is real. The instrument works. But the signal does not self-annotate. A mass spectrometer does not print "this is a protein." It prints peaks. The model says the peaks are a protein. Those are two different acts: the instrument producing data, and the model assigning meaning to that data. The rubric refuses to collapse them. When a paper says "we detected the protein," the rubric asks: what did you actually detect? A peak. What told you it was a protein? The model. Then the detection is not independent. The instrument is real. The interpretation is the model talking.
The replication rule refines this. Replication means another team ran the same model-dependent instrument and got the same model-dependent signal. That is not independence. Independence means someone who rejects the model could still see the thing. If a tumor shrinks, anyone can see it. If a fluorescence curve rises, only someone trained in the model can tell you what that means. Replication within the model is the model confirming itself at higher volume.
Test 2: Falsifiable
Root: Popper's demarcation criterion, extended to close the escape route Popper left open.
Popper said a claim is scientific if it can be falsified. If no observation could prove it wrong, it is not science. This is correct as far as it goes. What Popper did not fully address, and what Lakatos identified later, is that research programs survive Popper's criterion by adding protective belts. The core claim is never tested. When a prediction fails, the failure is absorbed by an auxiliary hypothesis: the equipment was faulty, the conditions were suboptimal, there were confounders, the variant was different, immunity had waned. The core claim is never at risk. The protective belt absorbs every negative result, and the program continues calling itself scientific because it reports p-values and confidence intervals.
The rubric closes this escape route. Test 2 does not ask whether the study reports a null hypothesis. It asks whether a null result could actually appear in the data, be visible, and be accepted by the field as a disproof. If the answer is no, because the groups were defined by model-dependent tests, because the outcomes were classified by model-dependent categories, because the field has absorbers ready for any negative, then the falsification is theatrical. The null exists on paper. The structure ensures it will never materialize.
This is the sophistry principle. A degenerating research program survives by performing the rituals of science. It states hypotheses. It reports statistics. It publishes in peer-reviewed journals. From the outside, it looks legitimate. The sleight of hand is in the design. The study is built so that the claim cannot die regardless of what the data shows. The null is offered as proof that testing occurred, but the groups, the outcome classifications, and the publication pipeline are arranged so that a null result is practically impossible to produce, report, or have accepted. The study looks falsifiable. It is not. The result was determined before the data was collected.
The structural vs. theatrical distinction is the heart of Test 2. Structural falsification is built into the experiment. All-cause death in two randomized groups is structural. If more people die in the treated arm, the result is visible, it cannot be reclassified, and the field cannot call it "confounding" without abandoning arithmetic. Theatrical falsification is a hypothetical condition the design ensures will never be met. A meta-analysis of observational studies inside the cholesterol hypothesis can report "we could have found no association," but every study inside the review used model-dependent group definitions, the field has infinite absorbers for a null result, and negative studies are not published. The null is theater. The result was predetermined.
The explanation rule refines this. Every scientist who reports a failure offers a reason. Pasteur explained why some flasks spoiled. Fleming explained why some colonies were not cleared. That is not the same as a paradigm absorber. An explanation addresses a specific failure. An absorber protects the paradigm from ever dying. "Asymptomatic carriers" is an absorber because it would be invoked regardless of the data. "This flask was improperly sealed" is an explanation because it addresses a specific case and does not protect the core claim from death. If every flask had spoiled, Pasteur's claim would have died. If every patient is a non-responder and the field still says "the treatment works, these are just complex cases," the field is using an absorber.
Test 3: Non-Circular
Root: the logical prohibition against a proof containing its conclusion as a premise.
This is Aristotle. It is the oldest principle in logic. If A proves B, and B is defined by A, you have not proven B. You have restated A. The proof is circular. The conclusion is the premise wearing different clothing.
The rubric applies this to scientific instruments. If the primers were designed from a sequence the model already identified as viral, the PCR assay is circular. It was built to find what the model named. Finding it proves nothing. The tool was constructed to produce a positive signal for the thing the model says exists. When it produces that signal, the model says: see, the thing exists. This is a closed loop. The instrument is the conclusion pretending to be evidence.
The control rule follows. "No template" is not a control for circularity. It is a control for contamination. It tells you whether the signal came from the sample or from the lab. It does not tell you whether the signal means what the model says it means. A genuine control would challenge the premise: run the assay on samples the model says should not contain the target, using a detection method that does not depend on the model's definitions. If the assay is positive on samples the model says are negative, the model has a problem. But most studies do not run this control, because the assay was designed within the model and the researchers do not think to challenge the model with its own tool.
The chain rule is critical. If the first link in a chain is model-dependent, the entire chain is model-dependent. You cannot score the links separately. If a patient cohort was sorted into a trial by PCR, and the trial's outcome was all-cause mortality measured by body count, the outcome measurement is non-circular but the sorting is circular. The trial inherits the circularity of the sorting. You cannot say "the outcome measurement passed, so the trial passes." The trial is only as independent as its most model-dependent step. If the first step presupposes the model, everything downstream is contaminated, regardless of how clean the later steps look.
PART IV: THE LOGICAL PROOF
Theorem: These three tests are necessary and sufficient for distinguishing a claim anchored to reality from a claim anchored to its own framework.
Proof of Necessity
A claim anchored to reality must meet three conditions:
(1) The observation must be independent of the model.
If the observation requires the model to be visible, the claim is not anchored to reality. It is anchored to the model. The model produces a signal, interprets the signal, and reports the interpretation as observation. Someone outside the model cannot see the claim. They can only see the signal. The claim exists only as the model's narration of its own output. Without Test 1, this passes undetected. A paper reports "we detected viral RNA" and the reader accepts it as observation. But the observation was a fluorescence curve. The viral RNA was the model's interpretation. Test 1 catches this by asking whether the observation survives the removal of the model. If it does not, the claim is ungrounded.
Therefore Test 1 is necessary. Remove it and any model can generate its own evidence by interpreting its own instruments. The claim becomes unfalsifiable in practice because the only people who can see the evidence are people who already accept the model.
(2) The claim must be able to die.
If no result can disprove the claim, the claim is not a description of reality. It is a tautology that accommodates any data. A claim that survives every possible observation is not a claim about the world. It is a claim about nothing. It has no content because it excludes nothing. "X causes Y, and if Y does not appear, X still causes Y but other factors interfered" is not a falsifiable claim. It is a statement that will be true regardless of what happens, and a statement that is true regardless of what happens is indistinguishable from a statement that is false. It carries no information.
But the falsifiability must be structural, not theatrical. A study can report a null hypothesis while designing its groups, outcomes, and publication pipeline so that the null cannot materialize. The null is offered as proof that testing occurred. The structure ensures the test can never fail. This is sophistry wearing the costume of science. Test 2 catches this by asking whether a disproof could actually appear in the data as collected, be visible, and be accepted by the field. If absorbers are ready, if the groups are model-defined, if null results are not published, the falsification is theater.
Therefore Test 2 is necessary. Remove it and any claim can survive indefinitely by absorbing every negative result into an auxiliary hypothesis. The claim becomes immortal, and an immortal claim is not science. It is dogma with statistics.
(3) The instrument must not have been built from the conclusion.
If the tool that detects the thing was designed from the model that defines the thing, the detection is circular. The tool was constructed to find what the model named. Finding it confirms the model. Not finding it means the tool failed. Either way, the model wins. The instrument is not testing the claim. It is executing the claim. The proof is the premise. The evidence is the conclusion. The loop is closed.
This is the most fundamental violation. A circular proof is not a proof at all. It is a restatement. "We designed primers from the sequence we believe is viral. The primers produced a signal. Therefore the sequence is viral." The primers were built from the conclusion. The signal confirms the primers worked. It does not confirm that the sequence is viral. The model told you it was viral. You built a tool to find it. The tool found it. The model says: confirmed. But the confirmation is the model confirming its own assumption through a tool that inherited the assumption.
Therefore Test 3 is necessary. Remove it and any model can generate its own proof by building instruments from its own definitions. The proof is a mirror. The model sees itself.
Proof of Necessity (Combined)
Remove Test 1: claims can be grounded in interpretation rather than observation, and nobody can tell the difference.
Remove Test 2: claims can be immortal, absorbing every negative, and nobody can kill them.
Remove Test 3: claims can be proven by instruments that were built to find what the claim already named, and the proof is a closed loop.
Each test closes a specific escape route for sophistry. All three are required because sophistry enters through any open route. A claim that passes Test 1 and Test 2 but fails Test 3 is observable and falsifiable but circularly proven. A claim that passes Test 2 and Test 3 but fails Test 1 is falsifiable and non-circular but based on model-dependent interpretation. A claim that passes Test 1 and Test 3 but fails Test 2 is independent and non-circular but cannot die. None of these is sufficient alone.
Proof of Sufficiency
If all three tests pass, the claim is anchored to reality through three independent chains:
- The observation exists without the model. Someone who rejects the entire framework can see the same thing. The claim is grounded in something external to the model.
- The claim can die. A result exists that would disprove it, and that result could actually appear in the data, be visible, and be accepted. The claim is at risk.
- The instrument was not built from the conclusion. The detection is independent of the definition. The proof is not a loop.
A claim meeting all three conditions cannot be sophistry. It is anchored to something external (Test 1), it risks death (Test 2), and it does not loop back on itself (Test 3). What remains after these three conditions is a claim that stands or falls on reality alone. That is what science claims to be. That is what these three tests enforce.
No fourth test is required because no fourth escape route exists. If the observation is independent, the claim can die, and the proof is not circular, there is no structural way for the claim to survive as sophistry. It might still be wrong. The observation might be misinterpreted, the causal mechanism might be different from what is claimed. But the evidence is structurally sound. The claim is scientific even if it is imperfect. Science does not require infallibility. It requires that the claim be grounded, at risk, and non-circular. These three tests enforce exactly those conditions.
No fourth condition can improve the filter because any fourth condition would either be redundant with one of the three (e.g., "the sample size must be adequate" is a quality concern, not a structural one; a small sample that passes all three tests is still scientific, just underpowered) or would impose a requirement that is not logically necessary (e.g., "the study must be replicated"; replication within the model is not independence, and replication of a non-circular study adds confidence but does not change the structural verdict). The three tests are complete.
QED
The three tests are necessary because each closes a specific escape route for sophistry. They are sufficient because closing all three routes leaves no structural path for a claim to survive as sophistry while passing all three. The filter is complete. It cannot be improved by addition. It can only be weakened by subtraction.
PART V: WHY IT CANNOT BE SANITIZED
This rubric cannot be co-opted, softened, or integrated into the existing apparatus of peer review without destroying what makes it work. Here is why:
It does not negotiate. Peer review negotiates. It asks "is this study good enough?" and accepts incremental quality on a spectrum. This rubric does not ask whether the study is good enough. It asks whether the study is structurally sound. There is no spectrum. The groups were either defined by model-dependent tests or they were not. The field either has absorbers or it does not. The instrument was either built from the conclusion or it was not. These are binary conditions. "Mostly non-circular" is not a score. "Somewhat falsifiable" is not a score. The rubric does not grant partial credit for sophisticated methodology that still fails the structural test.
It does not respect authority. A study from Harvard using PCR, sequencing, and antibody assays fails Test 1 and Test 3 regardless of the institution. A study from an unknown author counting visible outcomes in two groups passes. The rubric does not weight credentials, publication venue, or citation count. It weights logical structure. This is intolerable to the existing scientific apparatus, which is built on authority, consensus, and institutional hierarchy. Integrating this rubric into peer review would mean rejecting most published research in molecular biology, genetics, virology, and pharmacology. No journal will do that. No funding body will require it. The rubric cannot be sanitized because sanitizing it means making it acceptable to the system it exposes, and the system it exposes is the system that would have to accept it.
It names the fraud as structural, not individual. This is the most unsanitizable part. The rubric does not say "these scientists are lying." It says "the structure of this experiment cannot produce a disproof, the instrument is circular, and the observation is model-dependent." It attacks the methodology, not the methodologist. This means the scientists being scored cannot defend themselves by pointing to their integrity, their credentials, or their intent. The rubric does not care about intent. It cares about structure. A well-intentioned scientist running a circular assay has produced a circular result. Their good faith does not un-circularize it. Their expertise does not make the observation independent. Their peer review does not make the falsification structural. The rubric is immune to the defenses that scientists use to protect their work from methodological critique, because those defenses are all about the scientist, and the rubric is about the structure.
It exposes the paradigm, not just the study. This is what makes it dangerous at Layer 2. A study can pass all three tests and still operate inside a paradigm that fails all three. The dexamethasone trial counted deaths in two randomized groups. That is scientific methodology. The allopathic model that produced the trial has never been tested the same way. No study has ever taken the entire allopathic framework and asked: could any result kill it? The answer is no, because the framework has absorbers for every possible negative. "The drug was wrong, not the model." "The dose was wrong, not the model." "The patient population was wrong, not the model." The paradigm is immortal. An immortal paradigm is not science. It is dogma with clinical trials. The rubric says this explicitly. It scores the paradigm as unscientific even when individual studies within it are scientific. That distinction, that a scientifically conducted study can occur inside an unscientific framework, is the distinction that the scientific establishment cannot tolerate, because it means the establishment itself can be unscientific even when its individual products look rigorous.
It cannot be reframed as a quality metric. Quality metrics are negotiable. "This study has low power." "This study has selection bias." "This study needs better controls." These are improvements that can be made within the existing framework. The rubric does not offer improvements. It offers a verdict. The study is scientific, unscientific, mixed, or cannot tell. There is no "improve this and resubmit." If the study fails Test 3 because the instrument was built from the conclusion, the fix is not "use a better instrument." The fix is "build an instrument that does not presuppose the thing you are trying to prove." That fix requires abandoning the model's definitions, which means stepping outside the paradigm, which the paradigm cannot allow. The rubric's prescriptions are not implementable within the paradigm being tested. That makes it unsanitizable. You cannot integrate a test whose prescription is "leave the framework."
PART VI: THE LINE
This is the line. On one side: claims anchored to reality through independent observation, genuine risk of death, and non-circular proof. On the other side: everything else. Everything else includes most of published science. Everything else includes Nobel-winning work that used circular instruments. Everything else includes consensus positions endorsed by every major institution. The rubric does not care. The line is structural. It does not move for authority, consensus, or volume of publications.
Science is the only reliable method humans have for determining what is true about reality. The rubric measures whether a claim earned its connection to reality through sound method or assumed that connection and disguised the assumption as investigation. This measurement happens twice when a paradigm is present, once when it is not. When a paradigm exists, the paradigm verdict takes precedence. When no theoretical framework is present, the study verdict stands alone.
When a paradigm is present, the study is scored on whether its specific methods tested its specific claim honestly. The paradigm is scored on whether the framework it sits inside can ever be wrong. Both verdicts are real. But they are not equal. The paradigm verdict takes precedence because the framework determines what questions get asked, what counts as a result, and what failures get absorbed. A study can be methodologically clean, honest arithmetic, visible outcomes, genuine randomization, and still produce nothing of epistemic value if it sits inside a paradigm that has arranged things so it can never be challenged. The labor was real. The container decided what the labor was for. If the paradigm cannot die, nothing inside it is genuinely at risk, and a claim that is not at risk is not connected to reality regardless of how clean its methodology looks. The study verdict tells you whether the experiment was honest. The paradigm verdict tells you whether honesty was possible.
That is the rubric. It is three tests. It is two layers. The verdicts are mechanical. The logic is closed. The line does not move.