Abstract
A learning curve shows how fast a learning machine improves its behavior as the number of training examples increases. This paper proves a universal asymptotic behavior of learning curves for general noiseless dichotomy machines, or neural networks. It is proved that irrespective of the architecture of a machine, the average predictive entropy or the information gain 〈e*(t)〉 converges to 0 as 〈e*(t)〉 ∼ d/t as the number t of training exampies increases, where d is the number of modifiable parameters of a machine.
| Original language | English |
|---|---|
| Pages (from-to) | 161-166 |
| Number of pages | 6 |
| Journal | Neural Networks |
| Volume | 6 |
| Issue number | 2 |
| DOIs | |
| State | Published - 1993 |
| Externally published | Yes |
Keywords
- Entropic error
- Generalization error
- Information gain
- Learning curve
- Universal theorem
Fingerprint
Dive into the research topics of 'A universal theorem on learning curves'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver