THE EFFECT OF TEST LENGTH, SAMPLE SIZE, AND ITEM RESPONSE THEORY MODELS ON GENERALIZABILITY THEORY

Priarti Megawanti, Siti Alifah, Sulastry Pardede, Erna Megawati, Wardani Rahayu

Abstract


Purpose – This study examined the effects of test length, sample size, and Item Response Theory (IRT) models on the reliability of Generalizability Theory (GT) coefficients.

Methodology – Simulated dichotomous and polytomous data were generated using WinGen and analyzed using EduG. The study employed several combinations of sample sizes (N = 200, 300, 500, 1000, and 5000), test lengths (n = 5, 10, 30, 50, and 100), dichotomous models (1PL, 2PL, and 3PL), and polytomous models (GRM and GPCM) with different response categories.

Findings – The findings indicate that test length had a stronger influence on G-coefficient reliability than sample size. In dichotomous data, the 1PL and 2PL models generally produced higher G-coefficients than the 3PL model, especially with a test length of 50 items. In polytomous data, the GPCM model produced higher and more stable G-coefficients than the GRM model, particularly when the number of response categories and test length increased. The results indicate that optimal combinations of test length, sample size, parametric model, and response categories are important for achieving high measurement reliability. This study also demonstrates the effectiveness of Generalizability Theory and EduG software in evaluating reliability for both dichotomous and polytomous data.

Contribution – This research helps researchers determine appropriate model parameters, sample size, and test length. Furthermore, it can strengthen theory in empirical research

Keywords


Generalizability Theory; Item Response Theory; Test Length; Sample Size; EduG

Full Text:

PDF

References


Boone, W. J., & Noltemeyer, A. (2017). Rasch analysis: A primer for school psychology researchers and practitioners. Cogent Education, 4(1). https://doi.org/10.1080/2331186X.2017.1416898

Brennan, R. L. (2001). Generalizability theory: Statistics for social science and public policy. In New York: Springer-Verlag. (Vol. 30).

Brennan, R. L. (2011). Generalizability theory and classical test theory. Applied Measurement in Education, 24(1), 1–21. https://doi.org/10.1080/08957347.2011.532417

Brennan, R. L., & Liao, J. R. (2020). Generalizability Theory References: The First Sixty Years (Issue 53).

Cardinet, J., Johnson, S., & Pini, G. (2010). Applying Generalizability Theory using EduG. Routledge. https://doi.org/10.4324/9780203866948

Clauser, B. E. (2008). A Review of the EDUG Software for Generalizability Analysis. International Journal of Testing, 8(3), 296–301. https://doi.org/10.1080/15305050802262357

Cronbach, L. J., Gleser, G. C., Nanda, H., & Rajaratnam, N. (1972). The Dependability of Behavioral Measurements: Theory of Generalizability for Scores and Profiles. John Wiley & Sons, Inc. https://doi.org/10.3102/00028312011001054

Dai, S., Vo, T. T., Kehinde, O. J., He, H., & Xue, Y. (2021). Performance of Polytomous IRT Models With Rating Scale Data : An Investigation Over Sample Size, Instrument Length, and Missing Data. Frontiers in Education, 6(September), 1–18. https://doi.org/10.3389/feduc.2021.721963

Eason, S. (1989). Why Generalizability Theory Yields Better Results than Classical Test Theory. The Annual Meeting of the Mid-South Educational Research Association, Little Rock.

Fiangga, S., & Sari, Y. M. (2017). Analisis Generalisabilitas Multifaset pada Instrumen Penalaran Matematika SMP. Jurnal Elemen, 3(2), 118. https://doi.org/10.29408/jel.v3i2.398

Fitriyah, I. M. (2023). Pengembangan Soal AKM Numerasi Berbasis Komputer untuk Kelas XI SMA. Universitas Negeri Yogyakarta.

Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of Item Response Theory (D. S. Foster (ed.)). Sage Publication.

Han, K. T., & Amherst, M. (2007). WinGen: Windows Software That Generates Item Response Theory Parameters and Item Responses. Applied Psychological Measurement, 31(5), 457–459. https://doi.org/10.1177/0146621607299271

Handayani, R., Rahariyoso, D., Falani, I., Oktadinata, A., Daya, W. J., & Ockta, Y. (2024). Analyzing The Reliability of Thesis Assessment Instruments: Generalizability Theory. Journal of Education, Teaching, and Learning, 9(1), 47–55. https://doi.org/10.26737/jetl.v9i1.5715

Ikeda, T. (2026). A Monte Carlo simulation study of sample size requirements for the Graded Response Model. PLOS ONE, 21(4), 1–15. https://doi.org/10.1371/journal.pone.0347684

Irianty, R., Selang, R. A. R., Ihsan, H., Supriyati, Y., & Falani, I. (2025). Generalizability Theory Analysis of Local Curriculum Validation Instruments : Item Effects and Rater Consistency. Al-Ishlah: Jurnal Pendidikan, 17(4), 6963–6970. https://doi.org/10.35445/alishlah.v17i4.

Istiyono, E. (2016). The application of GPCM on the MMC test as a fair alternative assessment model in physics learning. Proceedings of the 3rd International Conference on Research, Implementation and Education of Mathematics and Science (ICRIEMS), May, 25–30.

Karami, H. (2012). The Relative Impact of Persons, Items, Subtests, and Academic Background on Performance on a Language Proficiency Test. Psychological Test and Assessment Modeling, 54(3), 211–226.

Kartono. (2008). Equating the combined dichotomous and polychotomous item test model in an achievement test. Jurnal Penelitian Dan Evaluasi Pendidikan, 12(2), 302–320.

Kyriazos, T. A. (2018). Applied Psychometrics: Sample Size and Sample Power Considerations in Factor Analysis (EFA, CFA) and SEM in General. Psychology, 9, 2207–2230. https://doi.org/10.4236/psych.2018.98126

Lakes, K. D., & Hoyt, W. T. (2013). Applications of Generalizability Theory to Clinical Child and Adolescent Psychology Research. J Clin Child Adolesc Psychol, 38(1), 144–165. https://doi.org/10.1080/15374410802575461

Masters, G. N. (1982). A Rasch Model for Partial Credit Scoring. Psychometrika, 47(2), 149–174. https://doi.org/10.1007/BF02296272

Nababan, D., Nurfahrunnisa, A., & Rachmaniar. (2025). Analisis Kualitas Butir Soal Tes Kemampuan Akademik (TKA) SMA di Kota Merauke. Journal of Authentic Research, 4(2), 2075–2083. https://doi.org/10.36312/ep2nry15

Ngozika, O. E. (2026). Determining Minimum Test Length for Reliable Assessment Using Generalizability Theory: Evidence from an Economics Multiple-choice Test. African Educational Research Journal, 14(1), 128–133. https://doi.org/ISSN: 2354-2160

Ostini, R., & Nering, M. L. (2006). Polytomous Item Response Theory Models. SAGE Publications, Inc.

Parriott, D. (2017). Using generalizability theory to investigate sources of variance of the Autism Diagnostic Observation Schedule-2 with trainees, in Dissertation Abstracts International Section A: Humanities and Social Sciences.

Reise, S. P., & Yu, J. (1990). Parameter Recovery in the Graded Response Model Using MULTILOG. Journal of Educational Measurement, 27(2), 133–144. https://doi.org/10.1111/j.1745-3984.1990.tb00738.x

Shavelson, R. J., & Webb, N. M. (1991). Generalizability Theory - A Primer.

Suliyanah, A., B. D., Jauhariyah, M. N. R., Misbah, Mahtari, S., Saregar, A., & Deta, U. A. (2021). A bibliometric analysis of minimum competency assessment research with VOS viewer related to the impact on physics education in 2019-2020. Journal of Physics: Conference Series, 2110(1). https://doi.org/10.1088/1742-6596/2110/1/012022

Susongko, P. (2016). Validation of science achievement test with the Rasch model. Jurnal Pendidikan IPA Indonesia, 5(2), 268–277. https://doi.org/10.15294/jpii.v5i2.7690

Susongko, P. (2015). Studi generalisabilitas tes tipe dua facet dengan menggunakan analisis varian tiga jalur. Prosiding Seminar Nasional Pendidikan, 1–9.

Teker, G. T., Guler, N., & Uyanik, G. K. (2015). Comparing The Effectiveness of SPSS and EduG Using Different Designs for Generalizability Theory. Kuram ve Uygulamada Eğitim Bilimleri, 15(3), 635–645. https://doi.org/10.12738/estp.2015.3.2278

White, M. C. (2017). Generalizability of Scores from Classroom Observation Instruments. University of Michigan.




DOI: https://doi.org/10.36987/jes.v13i5.9442

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Priarti Megawanti

Creative Commons License
This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

Lisensi Creative Commons
Jurnal Eduscience (JES) by LPPM Universitas Labuhanbatu is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY - NC - SA 4.0)