Statistical manifolds

Machine-learning models do not merely occupy parameter space. They induce families of probability distributions with geometry of their own. This note asks what that geometry buys us—and when the word “manifold” clarifies rather than decorates an argument.

Questions this note must answer

  • How does the Fisher information metric arise from distinguishability?
  • Which properties are invariant under reparameterization?
  • What is the precise relationship between statistical and learned representation manifolds?
  • Can information geometry explain any empirical behavior that ordinary optimization language cannot?

This working outline is growing into a worked essay with derivations, counterexamples, and reproducible visualizations.