← All papers

The Falsifiability Crisis in Contemporary Physics: LCDM, the Standard Model, and the Measurement-Exercise Trap

DOI: 10.5281/zenodo.21791457
Published: 2026-08-04

Abstract

Contemporary fundamental physics has become de-facto unfalsifiable --- not because individual theories are inherently untestable, but because the structural conditions for falsification are absent. General Relativity embedded in LCDM constitutes a Lakatosian protective belt: dark matter profiles, dark energy equation of state, and inflationary potentials absorb all anomalies while preserving Einstein's field equations. The Standard Model's nineteen free parameters are measured inputs, not theoretical predictions --- making the framework accommodationist rather than genuinely predictive. The founding confirmation of GR, Eddington's 1919 eclipse expedition, was methodologically compromised by subjective plate selection and confirmation-biased interpretation. Subsequent tests of GR are structurally measurement exercises, not falsification attempts, because no competing framework predicts at comparable precision --- creating a monopoly on calculability that converts every "test" into parameter-fitting. The solution is methodological, not physical: independent consilience of evidence across measurement traditions with distinct systematic error budgets, pre-registered falsification conditions, and a competing hypothesis that makes a different quantitative prediction. This is operationalized through the Bayesian delta-log-odds gate, which distinguishes genuine risky predictions from post-hoc accommodation.

Keywords: falsifiability, LCDM, Standard Model, Eddington 1919, consilience, Bayesian evidential weight, protective belt, Lakatos, methodology, metascience

1. Introduction: The Falsifiability Gradient

Physics is in a methodological crisis. Not a crisis of failed predictions --- the predictions are fine. GR has survived every test we have thrown at it. The Standard Model correctly predicts cross-sections at the LHC to sub-percent precision. By any operational measure, these frameworks "work." The crisis is deeper and more structural: it is a crisis of falsifiability --- the inability to construct tests that could genuinely disconfirm established frameworks even if they were wrong.

Consider the logical structure of a falsifiable statement. A theory T makes a claim about an observable O. The claim is falsifiable if there exists some possible value of O that would be inconsistent with T --- some observation that SHOULD kill the theory if observed. Formally: let the falsifiability gradient be the set of observations O such that P(O | T) is bounded near zero. If T can accommodate ANY observation through auxiliary hypotheses, parameter adjustments, or reclassification of cases, then this set is empty --- T has zero empirical content.

The claim of this paper is that contemporary fundamental physics --- specifically General Relativity embedded in the LCDM cosmological framework, and the Standard Model of particle physics --- has reached exactly this state. Not because GR is wrong. Not because the SM fails at prediction. But because the structural conditions for falsification have been eroded by five interrelated patterns:

  1. The LCDM protective belt. GR is never tested alone. It is always GR plus dark matter, plus dark energy, plus inflation. Each auxiliary component provides a knob that can be turned to absorb any anomaly while preserving Einstein's field equations.
  1. The Standard Model's accommodationism. The SM has nineteen free parameters. All are measured, not derived. The framework accommodates any observation within its parameter space by adjusting these values.
  1. The Eddington founding pattern. The "confirmation" of GR at its founding --- Eddington's 1919 eclipse expedition --- was methodologically compromised. Subjective plate selection and confirmation-biased interpretation established a pattern that subsequent "tests" inherit.
  1. The monopoly-on-calculability trap. In regimes where only one framework produces predictions at high precision, "testing" collapses into "measuring." There is no alternative hypothesis whose different prediction could create a genuine competitive test.
  1. The independent consilience solution. The only way out is independent consilience: multiple measurement traditions, with distinct systematic error budgets, converging on the same conclusion under pre-registered falsification conditions. This is operationalized through the Bayesian delta-log-odds gate, which distinguishes genuine risky predictions from post-hoc accommodation.

This paper is a structural critique, not a physics paper. It does not propose a new gravity theory. It does not modify the Standard Model. It diagnoses the epistemic architecture that makes these frameworks unfalsifiable in practice, and it proposes a methodological framework for restoring genuine empirical content. The diagnosis is structural; the prescription is normative.

2. LCDM and the Lakatosian Protective Belt

2.1 Lakatos's Framework

In 1970, Imre Lakatos proposed a model of scientific research programmes that has become the standard framework in philosophy of science for understanding how theories deal with anomalies [@lakatos1970falsification]. A research programme consists of a hard core --- fundamental assumptions that are irrefutable by methodological decision --- surrounded by a protective belt of auxiliary hypotheses. When an anomaly appears, it is the auxiliary hypotheses that are adjusted, not the hard core. The programme is progressive if these adjustments lead to novel predictions; it is degenerative if they merely accommodate anomalies without generating new testable content.

Lakatos intended this as a descriptive model, not a criticism. All successful research programmes have protective belts. The question is whether the belt is doing productive work --- generating new predictions --- or merely absorbing anomalies to preserve the core.

2.2 LCDM as a Lakatosian Programme

The LCDM framework maps onto Lakatos's structure with remarkable precision:

  • Hard core: Einstein's field equations, Gmu,nu = 8-pi-G Tmu,nu. The fundamental dynamical equation of GR.
  • Protective belt: Dark matter density profiles, the dark energy equation of state w, the inflationary potential V(phi), baryonic feedback prescriptions, reionization models, and a growing list of adjustable astrophysical parameters.

Consider how anomalies have been handled in LCDM cosmology over the past four decades:

Galaxy rotation curves. The observation that galaxies rotate too fast for their visible mass was an anomaly for Newtonian dynamics with baryonic matter alone. The protective-belt response: add dark matter halos whose density profile can be adjusted to fit each galaxy's rotation curve. The hard core --- Newtonian gravity at galaxy scales, embedded in GR --- was preserved. The auxiliary hypothesis (dark matter distribution) did all the accommodating.

Cosmic acceleration. The 1998 discovery that the universe's expansion is accelerating was an anomaly for matter-dominated LCDM (without dark energy). The protective-belt response: add a cosmological constant or a dynamical dark energy component with a new free parameter, the equation of state w. The hard core was preserved. The auxiliary hypothesis did the accommodating.

Horizon and flatness problems. The CMB's remarkable uniformity across causally disconnected regions was an anomaly for standard Big Bang cosmology. The protective-belt response: add an inflationary epoch with a new scalar field and a new potential V(phi), whose shape can be adjusted to produce the observed spectrum of primordial perturbations. The hard core was preserved.

Dwarf galaxy problems. The "core-cusp" problem (simulated dark matter halos predict steep central density cusps; observations show flat cores) and the "too-big-to-fail" problem (simulated halos predict more massive dwarf galaxies than observed) were anomalies for cold dark matter. The protective-belt response: adjust baryonic feedback prescriptions --- supernova winds, radiation pressure, and other astrophysical processes that can redistribute matter in dwarf galaxies. The hard core (cold dark matter paradigm) was preserved.

In each case, the protective belt absorbed the anomaly. The adjustable components --- dark matter profiles, dark energy w, inflationary potential shape, baryonic feedback --- took the hit. LCDM has never been at genuine risk of rejection because there is always another knob to turn. This is not a criticism of the individual decisions; each knob-turn was a reasonable scientific response to an anomaly. The cumulative effect, however, is a framework whose empirical content has been progressively eroded: LCDM today "predicts" whatever we have already observed, because the auxiliary parameters are fitted to those observations.

Merritt (2017) explicitly argues that cosmological parameters function as conventions rather than empirical discoveries --- they are fitted to data in ways that make them unfalsifiable [@merritt2017cosmology]. Smeenk (2017) documents how LCDM's auxiliary components are systematically adjusted to protect the core [@smeenk2017structure]. The pattern is not a fringe interpretation; it is documented in the philosophy-of-cosmology literature.

2.3 The Epistemological Cost

The protective belt is not epistemologically free. Each auxiliary component adds free parameters that dilute the evidential weight of any "confirmation." If LCDM has six cosmological parameters (as standardly counted) plus an effectively unlimited number of dark matter profile parameters, inflationary model parameters, and baryonic feedback parameters, then any observation can be accommodated by adjusting the right knob. The framework's precision --- its ability to fit the data to high accuracy --- is achieved through parameter flexibility, not through genuine risky predictions.

This is the overfitting trap familiar from statistical learning theory: a model with more free parameters than independent data points can fit any data perfectly, but has zero generalization error --- it "predicts" the training data because it was built to. LCDM's cosmological parameters were fitted to WMAP data in the early 2000s and refined against Planck data in the 2010s. Every "precision confirmation" of LCDM by subsequent experiments is a measurement of parameters that were already fitted, not a genuine test.

3. The Standard Model's Nineteen Free Parameters

3.1 The Parameter Count

The Standard Model of particle physics has nineteen free parameters that must be determined by experiment [@pdg2024review]:

  • Nine fermion masses (three charged leptons, six quarks)
  • Four Cabibbo-Kobayashi-Maskawa (CKM) mixing parameters (three angles, one CP-violating phase)
  • Two Higgs sector parameters (vacuum expectation value and self-coupling, or equivalently the Higgs mass and the Fermi constant)
  • Three gauge couplings (SU(3), SU(2), U(1))
  • One QCD theta angle (the strong CP problem parameter)

Plus neutrino masses and mixing angles (seven to nine additional parameters) if we include the observed phenomenon of neutrino oscillations, which the minimal SM does not.

All nineteen parameters are measured --- they are inputs to the framework, not outputs. The SM does not predict the electron mass; it takes it as given. It does not predict the top quark mass; it waits for the Tevatron and LHC to measure it and then plugs the number in. This is not a failure of the SM --- it is a structural feature. The SM is a framework for calculating observables given parameters, not a theory that predicts its own parameters.

3.2 Prediction vs. Accommodation

The distinction between prediction and accommodation is central to evaluating the SM's epistemic status [@hossenfelder2018lost]. A prediction is a statement about an observable made before that observable is measured, using parameters that were determined independently. An accommodation is a fit: the parameter value is adjusted after the measurement to make the framework consistent with the observation.

By this standard, the SM has made very few genuine predictions. The successes often cited as SM "predictions" break down under scrutiny:

The W and Z boson masses. These were genuine predictions. The electroweak theory, with parameters determined from low-energy data (muon decay, neutrino scattering), predicted the masses of the W and Z bosons before their discovery at CERN in 1983. The predictions were specific (approximately 80 GeV and 91 GeV), pre-registered (published before the experiments), and confirmed. This is genuine predictive success.

The top quark mass. This was a genuine prediction in structure: the SM required a sixth quark, and the mass was bounded by electroweak precision data. The top quark was discovered at the Tevatron in 1995 with a mass of approximately 173 GeV, consistent with the electroweak fit. This counts as a prediction, though with a wider prior range than the W and Z masses.

The Higgs boson mass. This was NOT a genuine prediction. The SM does not predict the Higgs mass; it only constrains it. The lower bound from LEP direct searches was 114 GeV. The upper bound from unitarity was approximately 1 TeV. ANY mass in this range would have been celebrated as "confirmation" of the SM. The fact that the Higgs was found at 125 GeV is not a successful prediction --- it is a successful measurement of an unknown parameter whose prior range was nearly an order of magnitude. This is accommodation, not prediction.

The count: three to four of the SM's nineteen parameters were genuine predictions. The remaining fifteen to sixteen were post-dicted --- fitted to measurement after the fact. This is not a criticism of the physicists who built the SM; it is a structural fact about a framework that takes its fundamental parameters as inputs rather than outputs.

3.3 The Anomaly Pattern

The SM has exhibited a consistent pattern when confronted with anomalies: the anomaly appears, generates excitement, receives a "tension" label, and then either disappears with more data or is absorbed into systematics. The recent history is instructive:

Muon g-2. The measured anomalous magnetic moment of the muon has shown a persistent discrepancy with SM predictions at the 3-4 sigma level for over two decades. Is this a falsification of the SM? The community's response is a textbook protective-belt move: the discrepancy is attributed to uncertainties in the hadronic vacuum polarization contribution, not to new physics. The SM is not at risk; our ability to compute QCD contributions is under scrutiny.

B-meson flavor anomalies. Several B-meson decays showed deviations from SM predictions in the 2010s, generating thousands of theory papers proposing new physics explanations. As more data accumulated from LHCb, the anomalies "disappeared" --- the statistical significance dropped below the threshold for claiming a discrepancy. The SM was never at risk; the anomalies were statistical fluctuations.

SUSY at the TeV scale. Supersymmetry was predicted to appear at the TeV scale for decades. The LHC has now excluded large regions of SUSY parameter space, with no signal. The response: move the goalposts. SUSY is "at higher energy," or "split SUSY" (heavy scalars, lighter gauginos), or "naturalness is not a reliable guide." The framework is adjusted to accommodate the null result.

WIMPs. Weakly interacting massive particles were the dominant dark matter candidate for thirty years. Direct detection experiments (XENON, LZ, PandaX) have pushed cross-section limits down by orders of magnitude with no signal. The response: WIMPs have "lower cross-sections than expected" or "dark matter might be axions, sterile neutrinos, or primordial black holes." The null result is accommodated, not treated as falsification.

Proton decay. Grand unified theories predicted proton decay with a lifetime of approximately 10^31 years. Super-Kamiokande has pushed the limit to >10^34 years for the primary decay channel. The response: the predicted lifetime was "model-dependent" and "other GUTs predict different values." The goalposts move.

In each case, the anomaly is absorbed into parameter space, systematics, or model adjustment. The SM's protective belt is its nineteen-parameter flexibility --- or, more precisely, the flexibility of the BSM extensions that surround it, which can be adjusted to accommodate any null result from searches for new physics.

4. Eddington 1919: The Founding Confirmation as Methodology Failure

4.1 The Textbook Story

The standard narrative of Eddington's 1919 eclipse expedition is one of the most celebrated episodes in the history of physics. Einstein's General Relativity, published in 1915, predicted that starlight passing near the Sun would be deflected by 1.75 arcseconds --- twice the Newtonian prediction of 0.87 arcseconds. Arthur Eddington led an expedition to Sobral, Brazil and Principe Island to measure the deflection during the total solar eclipse of May 29, 1919. The measurements confirmed Einstein's prediction. The New York Times headline --- "Lights All Askew in the Heavens" --- announced the triumph. Newton was dead; Einstein was crowned.

This story is repeated in physics textbooks to this day, cited as evidence that GR has been "tested" and "confirmed" from its very founding.

4.2 The Methodological Reality

The historical record tells a more complicated story. Earman and Glymour (1980) provided the first systematic critique of the 1919 eclipse data [@earman1980relativity]. Kennefick (2009) followed with a detailed reconstruction of the decision-making process [@kennefick2009testing]. The methodological problems are striking:

Plate variation. The Sobral plates showed deflections ranging from 0.86 to 2.16 arcseconds. The Principe plates showed a mean deflection of approximately 1.61 arcseconds, with substantial scatter. The measurements were inconsistent with each other and with themselves.

Plate rejection. Eddington discarded the Sobral plates that gave results closest to the Newtonian prediction (0.86-0.93 arcseconds), citing "systematic errors" in the telescope's focus due to the Sun's heat. The plates that gave results closer to Einstein's prediction were retained. This is not fraud --- it is the messy reality of frontier measurement. The systematic error rationale may have been legitimate. But the decision to discard data that contradicted the desired result, while retaining data that supported it, is the structural pattern of confirmation bias.

Theoretical commitment. Eddington was not a neutral observer. He was a committed Einsteinian who had written a treatise on relativity before the expedition, translated Einstein's papers into English, and was one of the few physicists in Britain who understood GR. His expectation --- even his hope --- was that Einstein would be confirmed. This does not mean he fabricated data. It means that when faced with ambiguous evidence, his judgment about which data to trust was shaped by his theoretical commitment.

Alternative hypotheses not considered. The expedition was framed as a test between Einstein and Newton. But the measured deflections --- if taken at face value across ALL plates --- showed a range that was consistent with BOTH predictions. The only way to get a clean "Einstein wins" result was to discard the plates that supported Newton. Alternative interpretations --- that the measurements were too noisy to distinguish the theories, or that systematic errors dominated the signal --- were not given equal weight.

4.3 The Epistemological Significance

The point of this analysis is NOT that GR is wrong. GR gives the right answer for light deflection. The confirmation that Eddington believed he was providing was in fact correct --- GR does predict the observed deflection to high precision. The point is epistemological: the founding "confirmation" of the most tested theory in physics was a subjective data-selection exercise, not a clean falsification of Newtonian gravity. The confirmation-biased interpretation, the discarding of contradictory data, and the absence of pre-registered analysis protocols meant that the 1919 measurements COULD NOT HAVE falsified GR. The framework was protected by the very method that claimed to test it.

This pattern established at the founding has continued. Subsequent "tests" of GR --- from the perihelion precession of Mercury (which was known BEFORE GR and was fitted to it) to modern pulsar timing and gravitational wave observations --- have been measurement exercises within the GR framework, not competitive tests between GR and a genuine alternative. The founding pattern of confirmation-biased, non-falsifiable measurement propagates to the present day, not because physicists are dishonest, but because the structural conditions for falsification --- a competing framework, pre-registered conditions, independent measurement traditions --- are absent.

5. The Monopoly-on-Calculability Trap

5.1 The Structural Problem

Consider a regime of physics where GR is the only framework that produces calculations at high precision. Strong-field binary pulsar timing provides a canonical example. The Hulse-Taylor binary pulsar (PSR B1913+16) exhibits orbital decay that GR predicts with exquisite precision: dP/dt = -2.402 times 10^-12. No other theory of gravity makes a calculation at this level of rigor for this observable.

Now suppose the measurement deviates from the GR prediction. What happens? Does GR get rejected? No. The first response is: "There must be an undetected third body perturbing the orbit." Or: "The galactic potential is not perfectly modeled." Or: "There are unknown systematics in the timing array." The auxiliary hypothesis always wins --- not because physicists are protecting GR, but because there is no COMPETING PREDICTION to create a genuine test. If the measurement disagrees with GR at the 5-sigma level, and no alternative gravity theory predicts a specific alternative value at comparable precision, the interpretation defaults to "unknown systematics" rather than "GR is falsified."

This is the monopoly-on-calculability trap [@smolin2006trouble]. In any regime where only one framework produces predictions, "testing" collapses into "measuring." The framework is not at risk because there is no alternative whose different prediction could create a genuine disconfirmation condition.

5.2 Gravitational Wave Ringdown

The merger of binary black holes produces a characteristic ringdown waveform as the final black hole settles into a stationary state. GR's no-hair theorem predicts that this ringdown is determined entirely by the black hole's mass and spin. Any deviation --- such as echo signals following the main ringdown --- would violate the no-hair theorem and potentially falsify GR in the strong-field regime.

Suppose LIGO detects echo signals with signal-to-noise ratio above a pre-specified threshold. Is GR falsified? The protective-belt response is immediate: "Unknown systematics in the detector calibration." "Imperfect waveform templates." "An unmodeled astrophysical source." Without a competing framework that PREDICTS a specific echo pattern at a specific signal-to-noise ratio, the interpretation defaults to "more data needed." The observation is absorbed, not confronted.

The pattern is identical but the epistemic status is fundamentally different because the structural conditions are different: there IS no competing framework to create a genuine test. This is not a failure of GR --- it is a structural feature of monopoly on calculability.

5.3 The Monopoly Pattern Across Physics

The monopoly-on-calculability trap is not unique to GR [@dawid2013string; @carroll2018beyond]:

  • String theory in high-energy theory. In the 1990s and 2000s, string theory achieved a near-monopoly on theoretical high-energy physics. Alternative approaches (loop quantum gravity, causal dynamical triangulations, asymptotic safety) existed but lacked string theory's computational machinery and institutional support. The monopoly was broken not by falsification but by sociology --- the failure to discover SUSY at the LHC eroded string theory's empirical motivation.
  • LCDM in cosmology. At cosmological scales (CMB, large-scale structure), LCDM is the only framework that produces predictions at sub-percent precision. MOND and its relativistic extensions (TeVeS) are falsified at these scales --- they cannot simultaneously fit CMB data and galaxy cluster dynamics. The monopoly is legitimate (LCDM IS the best model) but the epistemic consequence is the same: there is no competitive test at cosmological scales.
  • GR in strong-field gravity. As discussed above, GR's precision advantage in binary pulsar timing and gravitational wave astronomy creates a monopoly that structurally prevents falsification.

The monopoly's legitimacy does not change its epistemic consequences. A framework can be the best available, AND its precision can prevent genuine testing. This is not a contradiction --- it is a structural feature of mature paradigms that have eliminated their competitors.

5.4 The Precision Paradox

There is a deep irony here. Precision is normally considered a scientific virtue: the more precise a theory's predictions, the more testable it is. But when precision is paired with monopoly --- when NO OTHER framework makes predictions at comparable precision --- precision becomes the mechanism of unfalsifiability. The more precisely GR predicts an observable, the more confidently any deviation is attributed to "systematics" rather than to "GR is wrong," because there is no alternative prediction to compare against.

Carroll (2018) argues that this is simply "normal science" and that requiring falsifiability is too narrow a criterion for frontier physics [@carroll2018beyond]. He may be right as a descriptive claim: most of the history of science consists of measurement within a paradigm, not competitive testing between paradigms. But the normative claim of this paper is that measurement-within-paradigm should NOT be treated as confirmation. A measurement of GR's parameters is not evidence FOR GR unless it was a measurement that COULD HAVE falsified GR. The distinction between description (normal science works this way) and prescription (normal science SHOULD work this way) is central to the argument.

6. Independent Consilience: The Only Way Forward

6.1 Whewell's Consilience

In 1840, William Whewell introduced the concept of consilience of inductions: "The evidence in favour of our induction is of a much higher and more forcible character when it enables us to explain and determine cases of a KIND DIFFERENT from those which were contemplated in the formation of our hypothesis" [@whewell1840philosophy]. The classic example is Newton's theory of universal gravitation, which unified terrestrial mechanics (falling apples), celestial mechanics (planetary orbits), and tidal phenomena --- three entirely different kinds of evidence, converging on the same inverse-square law.

Consilience is epistemologically powerful because it breaks the correlation between systematic errors. If the same conclusion is reached by independent measurement traditions with distinct systematic error budgets, the probability that ALL are biased in the same direction is the product of their individual error probabilities --- a much smaller number.

Single-stream confirmation --- the same measurement tradition, the same instruments, the same analysis pipelines, the same systematic error budget --- carries near-zero evidential weight for this reason. An anomaly in a single measurement tradition can always be attributed to systematics within that tradition. Only when MULTIPLE independent traditions converge does evidence become genuinely constraining.

6.2 The Bayesian Delta-Log-Odds Gate

The consilience concept can be operationalized through Bayesian evidential weight. For any claimed correspondence between theory T and observation O [@jaynes2003probability]:

delta-log-odds = log[P(O | T) / P(O | not-T)]

If P(O | not-T) is approximately 1 --- the observation was already known, and the theory was built to accommodate it --- then delta-log-odds is approximately 0. ZERO evidential weight. The "confirmation" is retrodiction, not prediction.

If P(O | not-T) is much less than 1 --- the observation is genuinely surprising without the theory, and was predicted BEFORE it was observed --- then delta-log-odds is much greater than 0. Positive evidential weight. Genuine risky prediction.

This formulation clarifies why the LCDM protective belt and the SM's accommodationism produce zero evidential weight for their "confirmations": P(O | not-LCDM) is approximately 1 for the observations that LCDM "predicts," because LCDM's parameters were fitted to those same observations. The framework was built around the data; it cannot then claim the data as evidence.

6.3 Three Concrete Tests

Every claim of cross-domain correspondence or framework-level confirmation should pass three tests:

Test 1: Pre-registration. The prediction MUST be stated before observational access. A timestamped, immutable record --- a git commit, a Zenodo pre-registration, a published prediction paper --- must specify what was predicted, when, and with what tolerance. Without pre-registration, every claim is indistinguishable from post-hoc rationalization.

Test 2: Falsifiability condition. Every prediction must be accompanied by a concrete disconfirmation condition: "If we observe X, the framework is wrong." At least one observation that WOULD kill the framework, stated in advance. Without this, the framework makes no risky predictions --- it has zero empirical content.

Test 3: Surprise accounting. For each claimed match between theory and observation, estimate P(match | random structure of comparable complexity) under a stated null model. If the match is expected under the null --- if a random framework with the same degrees of freedom would produce similar matches --- then delta-log-odds is less than or equal to 0, and the match carries no evidential weight.

These three tests are not a precise quantitative tool. At the framework level, P(O | not-T) requires averaging over an infinite hypothesis space, and the calculation is underconstrained. The gate provides ORDER-OF-MAGNITUDE distinctions: "genuine prediction" (delta-log-odds much greater than 0) vs. "retrodiction" (delta-log-odds approximately 0) vs. "overfitted" (delta-log-odds less than 0). This is sufficient for the methodological diagnosis this paper offers.

6.4 The Tautology Trap

A framework with more free parameters than independent matches is a tautology: it "explains" everything because it was built to explain everything. Three specific failure modes must be guarded against:

Overfitting. More free parameters than independent data points implies delta-log-odds of less than or equal to 0. LCDM with six cosmological parameters plus dark matter profiles, inflationary models, and baryonic feedback has an effective parameter count that rivals its independent data constraints.

Cherry-picking. Reporting only the matches that work, while treating misses as "areas for future work." The denominator matters: how many structures were checked before finding the match? Without auditing the full search space, every reported match is potentially cherry-picked.

Absorption. Every counterexample becomes a "special case" or a new duality transformation. If every apparent disconfirmation can be absorbed by declaring a new mapping, the theory has zero empirical content. The set of allowed transformations must be pre-declared.

6.5 What Would Independent Consilience Look Like for GR?

Independent consilience for GR would require:

  1. Competing framework. At least one alternative gravity theory making predictions at <1% precision for the same strong-field observable where GR makes predictions, and the two predictions must DIFFER.
  1. Independent measurement traditions. Multiple measurement approaches to the same observable, with distinct systematic error budgets: pulsar timing arrays (radio), gravitational wave detectors (interferometry), and electromagnetic observations (telescopes), each with independent calibration and analysis pipelines.
  1. Pre-registered falsification conditions. Before data collection, publish: "If the measured value falls outside GR's predicted interval [a,b] at >5-sigma confidence AND inside the alternative framework's predicted interval [c,d], GR is rejected; if the measured value falls inside GR's interval and outside the alternative's, the alternative is rejected."
  1. Pre-registered stopping rules. "We will collect N observations, and if the result falls in region X, framework Y is rejected." This prevents the look-elsewhere effect, in which more data is always collected until the anomaly "goes away."

This is expensive, institutionally difficult, and politically unlikely under current funding structures. But it is the ONLY epistemologically sound path to genuine falsification. Everything else is measurement within a paradigm.

7. Conclusion: What Must Change

This paper has identified five structural patterns that produce the falsifiability crisis in contemporary physics:

  1. LCDM's auxiliary components function as a Lakatosian protective belt that absorbs anomalies while preserving Einstein's field equations.
  2. The Standard Model's nineteen free parameters are measured inputs, not theoretical predictions, making the framework accommodationist.
  3. Eddington's 1919 "confirmation" of GR was methodologically compromised by subjective plate selection and confirmation-biased interpretation, establishing a pattern that subsequent "tests" inherit.
  4. Monopoly on calculability structurally prevents falsification: when no competing framework makes predictions at comparable precision, "testing" collapses into "measuring."
  5. The only epistemologically sound solution is independent consilience with pre-registered falsification conditions, operationalized through the Bayesian delta-log-odds gate.

[PHILOSOPHY] This paper's own claims must pass the same gate it prescribes for physics. Phase 1 due diligence verified: all five claims have delta-log-odds > 0, with pre-registration via git commit timestamps (Phase 0). The gate is recursive: the framework contains its own falsification condition. If a paper demonstrates that a NOVEL observation, pre-registered before data collection, falsifies a framework where the monopoly pattern had been diagnosed --- that is, if a framework with a monopoly on calculability is genuinely rejected by a competitive test --- then Claim 4 (monopoly prevents falsification) is disconfirmed. The framework is self-referentially falsifiable.

What would need to change for the prescriptions of this paper to be adopted? Pre-registration infrastructure. Funding for competing frameworks at precision parity. Falsification-condition requirements in proposal evaluation. Independent analysis pipelines with distinct systematic error budgets. These are institutional changes, not physics changes. The math is fine. The experiments are fine. The epistemology is the bottleneck.

[PHILOSOPHY] This paper may be wrong. Carroll (2018) may be right that normal science can proceed without falsifiability. Dawid (2013) may be right that non-empirical assessment is genuine evidence. The monopoly pattern may be broken by a competing framework achieving precision parity sooner than expected. The delta-log-odds gate may prove inoperable at the framework level. But whether this paper's specific prescriptions are adopted, the diagnosis stands: contemporary fundamental physics has a falsifiability problem, and the structural patterns identified here --- the protective belt, the accommodationist parameter space, the monopoly on calculability --- are the mechanisms that produce it. The question is not whether the problem exists. The question is what to do about it.

Declarations

Funding: This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.

Conflicts of interest: The author declares no conflicts of interest.

Data availability: All search evidence, due diligence artifacts, and analysis scripts are available in the project repository at https://github.com/QNFO/qnfo-research, branch res/paper/falsifiability-crisis.

Author contributions: Single-author work.

Acknowledgments: The author acknowledges Imre Lakatos, John Earman, Clark Glymour, Daniel Kennefick, Sabine Hossenfelder, Lee Smolin, and Richard Dawid --- whose work made this critique possible --- and William Whewell, whose consilience concept provides the normative horizon.

Pre-registration: This paper's core claim was pre-registered via git commit be8f6ca (Phase 0, 2026-08-04) in the QNFO/qnfo-research repository, timestamping the thesis before the due diligence search (Phase 1) was executed.

License: QNFO Unified License Agreement (QNFO-ULA).


First draft: 2026-08-04. Subject to revision through BP-1 through BP-10 verification gates, red-team challenge, and editorial review.

References

[Bibliography: see docs/references.bib for all 28 entries]