← Back to Claims
MechanismUnknown — conflicting evidence

Implicit regularization by stochastic gradient descent is the primary explanation for why over-parameterized neural networks generalize

Weak-link patterns

Brief §8's named failure modes, computed from this claim's own paper roles and status -- not a single weakness score. Patterns B and E aren't implemented yet (need dependency-graph machinery, docs/BASIC_ROADMAP.md Phase 11).

C: Conflicting evidence

stored status is 'unknown_conflicting'

Open questions

Pattern G. Human-written, never generated -- deliberately no importance or priority score (docs/BASIC_ROADMAP.md's own precedent against inventing one, matching innovation_considerations).

Is SGD's implicit regularization the primary mechanism behind generalization in over-parameterized networks, or one of several redundant mechanisms (architecture, data augmentation, explicit regularization) that jointly suffice?

Motivated by: Pattern C (conflicting evidence -- direct contradiction plus a conceptual-replication failure)

Evidence map

Position shows direction, circle area shows citation count. Click a paper for the reason its role was assigned.

MECHANISMCLAIMUnknown — conflicting…Conceptual replication — success On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima 571 citationsKeskar 2016571 citesOriginal In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning 135 citationsNeyshabur 2014135 citesTheoretical support The Implicit Bias of Gradient Descent on Separable Data 15 citationsSoudry 201715 citesContradiction Sharp Minima Can Generalize For Deep Nets 147 citationsDinh 2017147 citesConceptual replication — failure Fantastic Generalization Measures and Where to Find Them 84 citationsJiang 202084 cites
SupportsDisconfirmsSynthesisCircle area ∝ citation count

Evidence dimensions

Eight independent 0–5 judgments, each with a written rationale. Deliberately never summed into one score.

Evidence strength

Real theoretical results exist for restricted settings (separable data, linear models), but nothing establishes implicit SGD regularization as THE primary explanation in the deep non-convex case the claim is about.

Replication (any kind)

Sub-results are examined repeatedly, but there is no single canonical experiment to reproduce.

Independent replication

No same-paradigm independent confirmation that this mechanism is the primary one; supporting work explores different facets rather than re-testing one result.

Methodological triangulationscored retroactively

Theory (margin/implicit-bias proofs), optimization experiments (batch size, minima sharpness) and large-scale measure comparisons have all been brought to bear - and they disagree.

Methodological quality

The theory is rigorous but proved under assumptions far from practice; the empirical work is solid but correlational.

Predictive success

The account's central promise - a generalization measure that predicts held-out performance - is what Jiang et al. tested directly and largely did not find.

Consensus

No consensus. Implicit bias, flat minima, margin, norm-based capacity and NTK/lazy-training accounts all remain live.

Alternative explanations ruled out

Many well-articulated competing mechanisms, none ruled out. The lowest score on this dimension in the corpus, and appropriately so.

Scored by human.

Why this status

The derivation rule, evaluated top to bottom, first match wins. This is the rule itself — not a confidence score standing in for one.

  1. 1
    Unknown — unexplored

    5 paper rows (needs <3) and evidence_strength=2 (needs <=1); both required

  2. 2
    Unknown — conflicting evidencematched

    (a) 1 failure vs 0 success rows -> no; (b) 0 review rows — opposite-verdict test is a human judgment, not a count; (c) 1 conceptual failure vs 1 conceptual success, 1 original -> fires (balanced)

  3. 3
    Unknown — underdeterminednot reached

    alternative_explanations=1 (needs <=1) AND evidence_strength=2 (needs >=3)

  4. 4
    Establishednot reached

    independent_replication=1 (>=4), evidence_strength=2 (>=4), consensus=2 (>=4)

  5. 5
    Strong, domain-limitednot reached

    evidence_strength=2 (>=4) AND independent_replication=1 (>=3), plus a written scope limit (human judgment)

  6. 6
    Active consensus, incompletenot reached

    consensus=2 (>=4) AND (methodological_quality=3 <=3 OR alternative_explanations=1 <=2)

  7. 7
    Plausible, under active investigationnot reached

    independent_replication=1 (needs ==2) AND evidence_strength=2 (needs 2-3), plus recent activity (human judgment)

  8. 8
    Speculativenot reached

    3/5 rows are empirical (needs 0) AND evidence_strength=2 (needs <=1)

  9. 9
    Preliminarymatchednot reached

    4 non-theoretical row(s) exist AND independent_replication=1 (needs <=1)

Paired claim

A status that belongs to a relationship between two claims, not to either one alone.

Phenomenon established, mechanism uncertain

observation is 'established' (established family) and mechanism is 'unknown_conflicting' (not settled). This is a property of the pair — both claims keep the status their own evidence earned, and neither was overwritten.

This claim is the mechanism · paired with the observation

Established

Heavily over-parameterized neural networks generalize well to held-out data despite having enough capacity to fit random labels

Why these are linked · reasonedPilot 5 (rubric v6). Claim B proposes implicit SGD regularization as the explanation for the phenomenon asserted by Claim A. Definitionally true of the two statements as written - B names A's effect as the thing it explains - so confidence_basis is 'reasoned' rather than 'empirical': no co-citation check can establish that one sentence is a proposed mechanism for another, which is a semantic relation between claims rather than a co-occurrence fact about terms (BASIC_METHODOLOGY Tier 1 §10 allows reasoned for definitionally-true claims, with a written rationale). This is the corpus's first claim-to-claim edge and the link that makes phenomenon_established_mechanism_uncertain derivable at all.

Recorded rationale

v5 rule 2 clause (c1): 1 conceptual_replication_failure (Jiang 2020) vs 1 conceptual_replication_success (Keskar 2016), within 1 with at least one of each. A contradiction row (Dinh 2017) exists but clause (a) needs one success-side row too, and there is none. Naive citation-only guess was 'preliminary' - the FIRST claim in the corpus not naively guessed 'established', and the first where the naive guess was too DISMISSIVE rather than too generous; see Pilot 5 log Findings 2 and 4 (OpenAlex fragments this field's arXiv/proceedings records, so counts measure indexing rather than influence). Pilot 5; prior familiarity disclosed HIGH.