Training, learning and inference: unified dynamics of neural systems
Abstract
Atomic generation facts compiled into a graph enable recursive scientific processes, and training-learning dynamics in nanoGPT are characterized by state-conditioned parameter updates, persistent functional reorganization, and frozen inference projections validated across architectures.
We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and relation role. Compiled into a Generation-Fact Graph (GFG), these facts provide an AI-native, compilable scientific fact substrate preserving generation histories. We establish a GFG-based recursive scientific process in which analysis, intervention, replay and validation form facts for later cycles. Using nanoGPT, we establish unified training-learning dynamics. Training is the evolution of a parameter-optimizer system with state and memory: each actual training action enters the receiving state and produces a finite-amplitude nonlinear functional response conditioned by that state and target-specific update geometry. Learning is the persistent reorganization of distributed functional support by these responses; capability formation, maintenance, decline or recovery becomes observable when target-specific states are evaluated against their readout boundaries. Three primary coordinates - target-boundary state, target-specific update geometry and parameter-Adam receiving state - yield a second-order predictor operating before post-update outputs are read. On held-out runs, it achieved 91.43% accuracy and 91.49% macro-averaged recall across four transitions. We further establish inference as a frozen projection of training-learning dynamics. Component gating and rollback show causal recruitment and non-additive combination of query-conditioned support formed during training, deriving organizational conditions realized by Attention. Controlled feedback indicates possible double-edged reinforcement effects. ResNet/CIFAR-100 and diffusion/CIFAR-10 experiments confirm receiving-state-conditioned responses, persistent support reorganization and frozen inference projection beyond nanoGPT.
Community
This work presents an experiment-first causal study of training, learning, inference, and feedback in realized neural networks. Across Transformer/Adam, ResNet/SGD momentum, and diffusion U-Net/AdamW systems, the same core relation structure is preserved: realized updates produce receiving-state-dependent nonlinear responses, learning persistently reorganizes distributed functional support, frozen inference recruits training-formed support, and feedback alters subsequent formation states. Only after completing these experiments and cross-system tests did I realize that I had experimentally demonstrated that neural networks are stateful dynamical systems. All protocols, experiments, independent checkers, and archived evidence are publicly available.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Delayed Optimizer-State Transport Shapes Short-Horizon Training Decisions (2026)
- Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control (2026)
- Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW (2026)
- Fiber Fingerprints of Hidden Learning-State Dynamics (2026)
- SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models (2026)
- HarnessWAM: Bridging Prediction and Deliberation in World Action Models (2026)
- A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper