← Back to Claims
ObservationEstablished

Heavily over-parameterized neural networks generalize well to held-out data despite having enough capacity to fit random labels

Weak-link patterns

Brief §8's named failure modes, computed from this claim's own paper roles and status -- not a single weakness score. Patterns B and E aren't implemented yet (need dependency-graph machinery, docs/BASIC_ROADMAP.md Phase 11).

No pattern fires on this claim.

Evidence map

Position shows direction, circle area shows citation count. Click a paper for the reason its role was assigned.

OBSERVATIONCLAIMEstablishedOriginal Understanding deep learning (still) requires rethinking generalization 2,361 citationsZhang 20212,361 citesConceptual replication — success Reconciling modern machine-learning practice and the classical bias–variance trade-off 1,559 citationsBelkin 20191,559 citesConceptual replication — success A closer look at memorization in deep networks 651 citationsArpit 2017651 citesConceptual replication — success Deep double descent: where bigger models and more data hurt* 147 citationsNakkiran 2021147 cites
SupportsDisconfirmsSynthesisCircle area ∝ citation count

Evidence dimensions

Eight independent 0–5 judgments, each with a written rationale. Deliberately never summed into one score.

Evidence strength

Directly demonstrated, trivially reproducible, and observed across essentially every architecture and dataset tried.

Replication (any kind)

Reproduced constantly, including as a standard teaching exercise - among the most-reproduced results in the corpus.

Independent replication

Confirmed by many groups with no connection to the original authors, in the same paradigm.

Methodological triangulationscored retroactively

Benchmark experiment, the double-descent risk-curve framing, and memorization-dynamics analysis all converge. Not 5 only because all of it is empirical; there is no independent theoretical derivation.

Methodological quality

Large, clean, easily-checked experiments; docked one point because ML rarely pre-registers and reporting standards are informal.

Predictive success

The implication that scaling capacity need not hurt generalization has held across a decade of much larger models.

Consensus

Not disputed by anyone.

Alternative explanations ruled out

Competing explanations concern why it happens, not whether; for this observation there is little left to explain away.

Scored by human.

Why this status

The derivation rule, evaluated top to bottom, first match wins. This is the rule itself — not a confidence score standing in for one.

  1. 1
    Unknown — unexplored

    4 paper rows (needs <3) and evidence_strength=5 (needs <=1); both required

  2. 2
    Unknown — conflicting evidence

    (a) 0 failure vs 0 success rows -> no; (b) 0 review rows — opposite-verdict test is a human judgment, not a count; (c) 0 conceptual failure vs 3 conceptual success, 1 original -> no

  3. 3
    Unknown — underdetermined

    alternative_explanations=4 (needs <=1) AND evidence_strength=5 (needs >=3)

  4. 4
    Establishedmatched

    independent_replication=5 (>=4), evidence_strength=5 (>=4), consensus=5 (>=4)

  5. 5
    Strong, domain-limitedmatchednot reached

    evidence_strength=5 (>=4) AND independent_replication=5 (>=3), plus a written scope limit (human judgment)

  6. 6
    Active consensus, incompletenot reached

    consensus=5 (>=4) AND (methodological_quality=4 <=3 OR alternative_explanations=4 <=2)

  7. 7
    Plausible, under active investigationnot reached

    independent_replication=5 (needs ==2) AND evidence_strength=5 (needs 2-3), plus recent activity (human judgment)

  8. 8
    Speculativenot reached

    3/4 rows are empirical (needs 0) AND evidence_strength=5 (needs <=1)

  9. 9
    Preliminarynot reached

    4 non-theoretical row(s) exist AND independent_replication=5 (needs <=1)

Paired claim

A status that belongs to a relationship between two claims, not to either one alone.

Phenomenon established, mechanism uncertain

observation is 'established' (established family) and mechanism is 'unknown_conflicting' (not settled). This is a property of the pair — both claims keep the status their own evidence earned, and neither was overwritten.

This claim is the observation · paired with the mechanism

Unknown — conflicting evidence

Implicit regularization by stochastic gradient descent is the primary explanation for why over-parameterized neural networks generalize

Why these are linked · reasonedPilot 5 (rubric v6). Claim B proposes implicit SGD regularization as the explanation for the phenomenon asserted by Claim A. Definitionally true of the two statements as written - B names A's effect as the thing it explains - so confidence_basis is 'reasoned' rather than 'empirical': no co-citation check can establish that one sentence is a proposed mechanism for another, which is a semantic relation between claims rather than a co-occurrence fact about terms (BASIC_METHODOLOGY Tier 1 §10 allows reasoned for definitionally-true claims, with a written rationale). This is the corpus's first claim-to-claim edge and the link that makes phenomenon_established_mechanism_uncertain derivable at all.

Recorded rationale

v5 rule 4: independent_replication=5, evidence_strength=5, consensus=5. Naive citation-only guess was 'established' - no divergence, which is the intended result: this is the corpus's SECOND true-negative control after Pilot 1's spacing effect. Pilot 5 (machine-learning theory; prior familiarity disclosed HIGH). Paired with the mechanism claim on implicit SGD regularization - see Pilot 5 log Finding 1: the pair should derive phenomenon_established_mechanism_uncertain, but that status is UNDERIVABLE because no schema mechanism can express a claim-to-claim link.