THE EFFECT OF TEST LENGTH, SAMPLE SIZE, AND ITEM RESPONSE THEORY MODELS ON GENERALIZABILITY THEORY
Abstract
Purpose – This study examined the effects of test length, sample size, and Item Response Theory (IRT) models on the reliability of Generalizability Theory (GT) coefficients.
Methodology – Simulated dichotomous and polytomous data were generated using WinGen and analyzed using EduG. The study employed several combinations of sample sizes (N = 200, 300, 500, 1000, and 5000), test lengths (n = 5, 10, 30, 50, and 100), dichotomous models (1PL, 2PL, and 3PL), and polytomous models (GRM and GPCM) with different response categories.
Findings – The findings indicate that test length had a stronger influence on G-coefficient reliability than sample size. In dichotomous data, the 1PL and 2PL models generally produced higher G-coefficients than the 3PL model, especially with a test length of 50 items. In polytomous data, the GPCM model produced higher and more stable G-coefficients than the GRM model, particularly when the number of response categories and test length increased. The results indicate that optimal combinations of test length, sample size, parametric model, and response categories are important for achieving high measurement reliability. This study also demonstrates the effectiveness of Generalizability Theory and EduG software in evaluating reliability for both dichotomous and polytomous data.
Contribution – This research helps researchers determine appropriate model parameters, sample size, and test length. Furthermore, it can strengthen theory in empirical researchKeywords
Full Text:
PDFReferences
Boone, W. J., & Noltemeyer, A. (2017). Rasch analysis: A primer for school psychology researchers and practitioners. Cogent Education, 4(1). https://doi.org/10.1080/2331186X.2017.1416898
Brennan, R. L. (2001). Generalizability theory: Statistics for social science and public policy. In New York: Springer-Verlag. (Vol. 30).
Brennan, R. L. (2011). Generalizability theory and classical test theory. Applied Measurement in Education, 24(1), 1–21. https://doi.org/10.1080/08957347.2011.532417
Brennan, R. L., & Liao, J. R. (2020). Generalizability Theory References: The First Sixty Years (Issue 53).
Cardinet, J., Johnson, S., & Pini, G. (2010). Applying Generalizability Theory using EduG. Routledge. https://doi.org/10.4324/9780203866948
Clauser, B. E. (2008). A Review of the EDUG Software for Generalizability Analysis. International Journal of Testing, 8(3), 296–301. https://doi.org/10.1080/15305050802262357
Cronbach, L. J., Gleser, G. C., Nanda, H., & Rajaratnam, N. (1972). The Dependability of Behavioral Measurements: Theory of Generalizability for Scores and Profiles. John Wiley & Sons, Inc. https://doi.org/10.3102/00028312011001054
Dai, S., Vo, T. T., Kehinde, O. J., He, H., & Xue, Y. (2021). Performance of Polytomous IRT Models With Rating Scale Data : An Investigation Over Sample Size, Instrument Length, and Missing Data. Frontiers in Education, 6(September), 1–18. https://doi.org/10.3389/feduc.2021.721963
Eason, S. (1989). Why Generalizability Theory Yields Better Results than Classical Test Theory. The Annual Meeting of the Mid-South Educational Research Association, Little Rock.
Fiangga, S., & Sari, Y. M. (2017). Analisis Generalisabilitas Multifaset pada Instrumen Penalaran Matematika SMP. Jurnal Elemen, 3(2), 118. https://doi.org/10.29408/jel.v3i2.398
Fitriyah, I. M. (2023). Pengembangan Soal AKM Numerasi Berbasis Komputer untuk Kelas XI SMA. Universitas Negeri Yogyakarta.
Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of Item Response Theory (D. S. Foster (ed.)). Sage Publication.
Han, K. T., & Amherst, M. (2007). WinGen: Windows Software That Generates Item Response Theory Parameters and Item Responses. Applied Psychological Measurement, 31(5), 457–459. https://doi.org/10.1177/0146621607299271
Handayani, R., Rahariyoso, D., Falani, I., Oktadinata, A., Daya, W. J., & Ockta, Y. (2024). Analyzing The Reliability of Thesis Assessment Instruments: Generalizability Theory. Journal of Education, Teaching, and Learning, 9(1), 47–55. https://doi.org/10.26737/jetl.v9i1.5715
Ikeda, T. (2026). A Monte Carlo simulation study of sample size requirements for the Graded Response Model. PLOS ONE, 21(4), 1–15. https://doi.org/10.1371/journal.pone.0347684
Irianty, R., Selang, R. A. R., Ihsan, H., Supriyati, Y., & Falani, I. (2025). Generalizability Theory Analysis of Local Curriculum Validation Instruments : Item Effects and Rater Consistency. Al-Ishlah: Jurnal Pendidikan, 17(4), 6963–6970. https://doi.org/10.35445/alishlah.v17i4.
Istiyono, E. (2016). The application of GPCM on the MMC test as a fair alternative assessment model in physics learning. Proceedings of the 3rd International Conference on Research, Implementation and Education of Mathematics and Science (ICRIEMS), May, 25–30.
Karami, H. (2012). The Relative Impact of Persons, Items, Subtests, and Academic Background on Performance on a Language Proficiency Test. Psychological Test and Assessment Modeling, 54(3), 211–226.
Kartono. (2008). Equating the combined dichotomous and polychotomous item test model in an achievement test. Jurnal Penelitian Dan Evaluasi Pendidikan, 12(2), 302–320.
Kyriazos, T. A. (2018). Applied Psychometrics: Sample Size and Sample Power Considerations in Factor Analysis (EFA, CFA) and SEM in General. Psychology, 9, 2207–2230. https://doi.org/10.4236/psych.2018.98126
Lakes, K. D., & Hoyt, W. T. (2013). Applications of Generalizability Theory to Clinical Child and Adolescent Psychology Research. J Clin Child Adolesc Psychol, 38(1), 144–165. https://doi.org/10.1080/15374410802575461
Masters, G. N. (1982). A Rasch Model for Partial Credit Scoring. Psychometrika, 47(2), 149–174. https://doi.org/10.1007/BF02296272
Nababan, D., Nurfahrunnisa, A., & Rachmaniar. (2025). Analisis Kualitas Butir Soal Tes Kemampuan Akademik (TKA) SMA di Kota Merauke. Journal of Authentic Research, 4(2), 2075–2083. https://doi.org/10.36312/ep2nry15
Ngozika, O. E. (2026). Determining Minimum Test Length for Reliable Assessment Using Generalizability Theory: Evidence from an Economics Multiple-choice Test. African Educational Research Journal, 14(1), 128–133. https://doi.org/ISSN: 2354-2160
Ostini, R., & Nering, M. L. (2006). Polytomous Item Response Theory Models. SAGE Publications, Inc.
Parriott, D. (2017). Using generalizability theory to investigate sources of variance of the Autism Diagnostic Observation Schedule-2 with trainees, in Dissertation Abstracts International Section A: Humanities and Social Sciences.
Reise, S. P., & Yu, J. (1990). Parameter Recovery in the Graded Response Model Using MULTILOG. Journal of Educational Measurement, 27(2), 133–144. https://doi.org/10.1111/j.1745-3984.1990.tb00738.x
Shavelson, R. J., & Webb, N. M. (1991). Generalizability Theory - A Primer.
Suliyanah, A., B. D., Jauhariyah, M. N. R., Misbah, Mahtari, S., Saregar, A., & Deta, U. A. (2021). A bibliometric analysis of minimum competency assessment research with VOS viewer related to the impact on physics education in 2019-2020. Journal of Physics: Conference Series, 2110(1). https://doi.org/10.1088/1742-6596/2110/1/012022
Susongko, P. (2016). Validation of science achievement test with the Rasch model. Jurnal Pendidikan IPA Indonesia, 5(2), 268–277. https://doi.org/10.15294/jpii.v5i2.7690
Susongko, P. (2015). Studi generalisabilitas tes tipe dua facet dengan menggunakan analisis varian tiga jalur. Prosiding Seminar Nasional Pendidikan, 1–9.
Teker, G. T., Guler, N., & Uyanik, G. K. (2015). Comparing The Effectiveness of SPSS and EduG Using Different Designs for Generalizability Theory. Kuram ve Uygulamada Eğitim Bilimleri, 15(3), 635–645. https://doi.org/10.12738/estp.2015.3.2278
White, M. C. (2017). Generalizability of Scores from Classroom Observation Instruments. University of Michigan.
DOI: https://doi.org/10.36987/jes.v13i5.9442
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Priarti Megawanti

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.








1.jpg)






