A Generalized Saturation Model and a Data-Saturation Theory for the Pricing of Artificial-Intelligence Training Data

Authors

  • Michael Willie Council for Medical Schemes

DOI:

https://doi.org/10.32493/sm.v8i2.60720

Keywords:

saturation model; nonlinear regression; diminishing returns; neural scaling laws; data pricing; Hill function; Michaelis–Menten; model selection; identifiability; artificial intelligence.

Abstract

The fees charged for access to data used to train and operate artificial-intelligence (AI) systems are widely observed to rise with the volume of data and then to flatten, a pattern this paper interprets as saturation. I propose a data-saturation theory of AI data pricing, grounded in the empirical regularity that model performance improves as a power law in data with diminishing returns, and I show that a rational, value-based fee schedule inherits this diminishing-returns structure and therefore saturates toward a ceiling. To represent the resulting curve I introduce the generalized data-saturation (GDS) model, a three-parameter saturation kernel that nests the Michaelis–Menten, Hill and stretched-exponential families as special cases and is governed by interpretable scale, sharpness and tail-curvature parameters. I develop the model’s analytic properties the half-saturation volume, the marginal-fee and elasticity functions, and a saturation index and a full inferential framework based on nonlinear least squares, including identifiability conditions, the Jacobian, asymptotic normality and model selection by information criteria. A simulation study (synthetic data) shows that the ceiling, baseline and scale parameters are recovered with low bias and near-nominal confidence-interval coverage, while the shape parameters are estimable but weakly identified under high noise or small samples, a property reported transparently. The framework gives analysts a principled, parsimonious tool for describing and comparing AI data-fee schedules, and a vocabulary for reasoning about where, and how sharply, the value of additional data saturates.

References

1. Akaike, H. (1974). A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19(6), 716–723.

2. Bates, D. M., & Watts, D. G. (1988). Nonlinear Regression Analysis and Its Applications. New York: Wiley.

3. Burnham, K. P., & Anderson, D. R. (2002). Model Selection and Multimodel Inference (2nd ed.). New York: Springer.

4. Davidian, M., & Giltinan, D. M. (1995). Nonlinear Models for Repeated Measurement Data. London: Chapman & Hall.

5. Draper, N. R., & Smith, H. (1998). Applied Regression Analysis (3rd ed.). New York: Wiley.

6. Efron, B., & Tibshirani, R. J. (1993). An Introduction to the Bootstrap. New York: Chapman & Hall.

7. Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., … Zhou, Y. (2017). Deep learning scaling is predictable, empirically. arXiv:1712.00409.

8. Hill, A. V. (1910). The possible effects of the aggregation of the molecules of haemoglobin on its dissociation curves. Journal of Physiology, 40(Suppl.), iv–vii.

9. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., … Sifre, L. (2022). Training compute-optimal large language models. arXiv:2203.15556.

10. Jennrich, R. I. (1969). Asymptotic properties of non-linear least squares estimators. Annals of Mathematical Statistics, 40(2), 633–643.

11. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., … Amodei, D. (2020). Scaling laws for neural language models. arXiv:2001.08361.

12. Michaelis, L., & Menten, M. L. (1913). Die Kinetik der Invertinwirkung. Biochemische Zeitschrift, 49, 333–369.

13. Pinheiro, J. C., & Bates, D. M. (2000). Mixed-Effects Models in S and S-PLUS. New York: Springer.

14. Ratkowsky, D. A. (1983). Nonlinear Regression Modeling: A Unified Practical Approach. New York: Marcel Dekker.

15. Schwarz, G. (1978). Estimating the dimension of a model. Annals of Statistics, 6(2), 461–464.

16. Seber, G. A. F., & Wild, C. J. (1989). Nonlinear Regression. New York: Wiley.

17. Sun, C., Shrivastava, A., Singh, S., & Gupta, A. (2017). Revisiting unreasonable effectiveness of data in deep learning era. Proceedings of the IEEE International Conference on Computer Vision (ICCV), 843–852.

18. Wu, C. F. J. (1981). Asymptotic theory of nonlinear least squares estimation. Annals of Statistics, 9(3), 501–513.

Downloads

Published

2026-08-31

How to Cite

Willie, M. (2026). A Generalized Saturation Model and a Data-Saturation Theory for the Pricing of Artificial-Intelligence Training Data. STATMAT: Jurnal Statistika Dan Matematika, 8(2), 378–387. https://doi.org/10.32493/sm.v8i2.60720

Issue

Section

Articles

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.