Classical test theory and item response theory validation of a building design and modeling fundamentals test in vocational education

Abstract

Classical test theory (CTT) and item response theory (IRT) rest on different measurement assumptions: CTT describes items relative to the tested group, whereas IRT places items and persons on a common latent ability scale. Using both frameworks therefore offers complementary validity evidence for classroom tests in vocational education, where such evidence is rarely reported. This research and development study developed a test for the Fundamentals of Building Design, Modeling, and Information subject (Dasar-dasar Desain Pemodelan dan Informasi Bangunan, Basic DPIB) and examined its items with CTT and the two-parameter logistic (2PL) IRT model. Forty multiple-choice items were written from a competency-based blueprint and judged by five experts (Aiken's V = 0.80–0.95). After a tryout with 30 students, 32 items were administered to 103 DPIB students selected by proportional stratified random sampling. CTT showed high split-half reliability (Spearman–Brown = 0.95), good-to-very-good discrimination for 30 items, and 13 items with at least one poorly functioning distractor; no item was classified as difficult. The 2PL model produced positive discrimination (a = 0.848 to 4.833) and mostly negative difficulty (b = −2.597 to 0.034), and test information peaked at θ ≈ −0.98. Both frameworks identified the same easiest, hardest, and most discriminating items, but they contributed different evidence: CTT pinpointed distractors to revise, whereas IRT showed that measurement precision is concentrated below average ability. The test is therefore better suited to identifying students who have not yet mastered basic competencies than to distinguishing high achievers; adding difficult items and recalibrating with a larger sample are required before high-stakes use.

Keywords
  • Test instrument
  • Vocational education
  • Classical test theory
  • Item response theory
  • 2PL
References
  1. Aiken, L. R. (1985). Three coefficients for analyzing the reliability and validity of ratings. Educational and Psychological Measurement, 45(1), 131–142. https://doi.org/10.1177/0013164485451012
  2. Ariansyah, K., Wismayanti, Y. F., Savitri, R., Listanto, V., Aswin, A., Ahad, M. P. Y., & Cahyarini, B. R. (2024). Comparing labor market performance of vocational and general school graduates in Indonesia: Insights from stable and crisis conditions. Empirical Research in Vocational Education and Training, 16, 5. https://doi.org/10.1186/s40461-024-00160-6
  3. Arieli-Attali, M., Ward, S., Thomas, J., Deonovic, B., & von Davier, A. A. (2019). The expanded evidence-centered design (e-ECD) for learning and assessment systems: A framework for incorporating learning goals and processes within assessment design. Frontiers in Psychology, 10, 853. https://doi.org/10.3389/fpsyg.2019.00853
  4. Awopeju, O. A., & Afolabi, E. R. I. (2016). Comparative analysis of classical test theory and item response theory based item parameter estimates of senior school certificate mathematics examination. European Scientific Journal, 12(28), 263. https://doi.org/10.19044/esj.2016.v12n28p263
  5. Ayanwale, M. A., Chere-Masopha, J., & Morena, M. C. (2022). The classical test or item response measurement theory: The status of the framework at the Examination Council of Lesotho. International Journal of Learning, Teaching and Educational Research, 21(8), 384–406. https://doi.org/10.26803/ijlter.21.8.22
  6. Bichi, A. A., & Talib, R. (2018). Item response theory: An introduction to latent trait models to test and item development. International Journal of Evaluation and Research in Education, 7(2), 142–151. https://doi.org/10.11591/ijere.v7i2.12900
  7. Boone, W. J. (2016). Rasch analysis for instrument development: Why, when, and how? CBE—Life Sciences Education, 15(4), rm4. https://doi.org/10.1187/cbe.16-04-0148
  8. Butakor, P. K. (2022). Using classical test and item response theories to evaluate psychometric quality of teacher-made test in Ghana. European Scientific Journal, 18(1), 139. https://doi.org/10.19044/esj.2022.v18n1p139
  9. Cook, D. A., Brydges, R., Ginsburg, S., & Hatala, R. (2015). A contemporary approach to validity arguments: A practical guide to Kane's framework. Medical Education, 49(6), 560–575. https://doi.org/10.1111/medu.12678
  10. Fan, X. (1998). Item response theory and classical test theory: An empirical comparison of their item/person statistics. Educational and Psychological Measurement, 58(3), 357–381. https://doi.org/10.1177/0013164498058003001
  11. Gierl, M. J., Bulut, O., Guo, Q., & Zhang, X. (2017). Developing, analyzing, and using distractors for multiple-choice tests in education: A comprehensive review. Review of Educational Research, 87(6), 1082–1116. https://doi.org/10.3102/0034654317726529
  12. Gyamfi, A., & Acquaye, R. (2023). Parameters and models of Item Response Theory (IRT): A review of literature. Acta Educationis Generalis, 13(3), 68–78. https://doi.org/10.2478/atd-2023-0022
  13. Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education, 15(3), 309–333. https://doi.org/10.1207/S15324818AME1503_5
  14. Istiyono, E., Dwandaru, W. S. B., Setiawan, R., & Megawati, I. (2020). Developing of computerized adaptive testing to measure physics higher order thinking skills of senior high school students and its feasibility of use. European Journal of Educational Research, 9(1), 91–101. https://doi.org/10.12973/eu-jer.9.1.91
  15. Iwintolu, R. O., Opesemowo, O. A. G., & Adetutu, P. O. (2024). Effect of 2-PL and 3-PL models on the ability estimate in mathematics binary items. Journal on Efficiency and Responsibility in Education and Science, 17(3), 257–272. https://doi.org/10.7160/eriesj.2024.170308
  16. Jabrayilov, R., Emons, W. H. M., & Sijtsma, K. (2016). Comparison of classical test theory and item response theory in individual change assessment. Applied Psychological Measurement, 40(8), 559–572. https://doi.org/10.1177/0146621616664046
  17. Kania, N., Kusumah, Y. S., Dahlan, J. A., Nurlaelah, E., Gurbuz, F., & Bonyah, E. (2024). Constructing and providing content validity evidence through the Aiken's V index based on the experts' judgments of the instrument to measure mathematical problem-solving skills. REID (Research and Evaluation in Education), 10(1), 64–79. https://doi.org/10.21831/reid.v10i1.71032
  18. König, C., Spoden, C., & Frey, A. (2020). An optimized Bayesian hierarchical two-parameter logistic model for small-sample item calibration. Applied Psychological Measurement, 44(4), 311–326. https://doi.org/10.1177/0146621619893786
  19. Kılıç, A. F., Koyuncu, İ., & Uysal, İ. (2023). Scale development based on Item Response Theory: A systematic review. International Journal of Psychology and Educational Studies, 10(1), 209–223. https://doi.org/10.52380/ijpes.2023.10.1.982
  20. Manurung, G. V., & Mulyana, R. (2024). Pengaruh penerapan metode AIR (Auditory, Intelektual, Repetisi) terhadap hasil belajar Dasar DPIB siswa SMK. Jurnal Pendidikan Teknik Bangunan, 4(2), 115–124. https://doi.org/10.17509/jptb.v4i2.75533
  21. Murphy, D. H., Little, J. L., & Bjork, E. L. (2023). The value of using tests in education as tools for learning—not just for assessment. Educational Psychology Review, 35, 89. https://doi.org/10.1007/s10648-023-09808-3
  22. Nabil, N. R. A., Wulandari, I., Yamtinah, S., Retno, S. D. A., & Ulfa, M. (2022). Analisis indeks Aiken untuk mengetahui validitas isi instrumen asesmen kompetensi minimum berbasis konteks sains kimia. PAEDAGOGIA, 25(2), 184–191. https://doi.org/10.20961/paedagogia.v25i2.64566
  23. Nisfatulsanah, A., & Sugiharto, B. (2024). Comparing classical test theory and item response theory for respiratory system question instruments. PIONIR: Jurnal Pendidikan, 13(3), 70–86. https://doi.org/10.22373/pjp.v13i3.25435
  24. Nurjanah, S., Iqbal, M., Zafrullah, Z., Mahmud, M. N., Seran, D. S. F., Suardi, I. K., & Arriza, L. (2024). Psychometric quality of multiple-choice tests under classical test theory (CTT): AnBuso, Iteman, and R. Jurnal Penelitian dan Evaluasi Pendidikan, 28(2), 161–172. https://doi.org/10.21831/pep.v28i2.71542
  25. Purnama, D. N. (2017). Characteristics and equation of accounting vocational theory trial test items for vocational high schools by subject-matter teachers' forum. REID (Research and Evaluation in Education), 3(2), 152–162. https://doi.org/10.21831/reid.v3i2.18121
  26. Purnama, D. N., & Alfarisa, F. (2020). Karakteristik butir soal try out teori kejuruan akuntansi SMK berdasarkan teori tes klasik dan teori respons butir. Jurnal Pendidikan Akuntansi Indonesia, 18(1), 36–46. https://doi.org/10.21831/jpai.v18i1.31457
  27. Rezigalla, A. A., Eleragi, A. M. E. S. A., Elhussein, A. B., Alfaifi, J., ALGhamdi, M. A., Al Ameer, A. Y., Yahia, A. I. O., Mohammed, O. A., & Adam, M. I. E. (2024). Item analysis: The impact of distractor efficiency on the difficulty index and discrimination power of multiple-choice items. BMC Medical Education, 24, 445. https://doi.org/10.1186/s12909-024-05433-y
  28. Sacks, R., & Pikas, E. (2013). Building information modeling education for construction engineering and management. I: Industry requirements, state of the art, and gap analysis. Journal of Construction Engineering and Management, 139(11), 04013016. https://doi.org/10.1061/(ASCE)CO.1943-7862.0000759
  29. Şahin, A., & Anıl, D. (2017). The effects of test length and sample size on item parameters in item response theory. Educational Sciences: Theory & Practice, 17(1), 321–335. https://doi.org/10.12738/estp.2017.1.0270
  30. Schroeders, U., & Gnambs, T. (2025). Sample-size planning in item-response theory: A tutorial. Advances in Methods and Practices in Psychological Science, 8(1). https://doi.org/10.1177/25152459251314798
  31. Setiawati, F. A., Amelia, R. N., Sumintono, B., & Purwanta, E. (2023). Study item parameters of classical and modern theory of differential aptitude test: Is it comparable? European Journal of Educational Research, 12(2), 1097–1107. https://doi.org/10.12973/eu-jer.12.2.1097
  32. Suharno, Pambudi, N. A., & Harjanto, B. (2020). Vocational education in Indonesia: History, development, opportunities, and challenges. Children and Youth Services Review, 115, 105092. https://doi.org/10.1016/j.childyouth.2020.105092
  33. Taber, K. S. (2018). The use of Cronbach's alpha when developing and reporting research instruments in science education. Research in Science Education, 48(6), 1273–1296. https://doi.org/10.1007/s11165-016-9602-2
  34. Tarrant, M., Ware, J., & Mohammed, A. M. (2009). An assessment of functioning and non-functioning distractors in multiple-choice questions: A descriptive analysis. BMC Medical Education, 9, 40. https://doi.org/10.1186/1472-6920-9-40
  35. Zanon, C., Hutz, C. S., Yoo, H., & Hambleton, R. K. (2016). An application of item response theory to psychological test development. Psicologia: Reflexão e Crítica, 29, 18. https://doi.org/10.1186/s41155-016-0040-x