CAT Field-Test Item Calibration Sample Size: How Large is Large under the Rasch Model?

Authors

  • Wei He

Keywords:

field-test item calibration, calibration sample size, computerized adaptive test, pretest item calibration, WINSTEPS

Abstract

This study was conducted in an attempt to provide guidelines for practitioners regarding the optimal minimum calibration sample size for pretest item estimation in the computerized adaptive test (CAT) under WINSTEPS when the fixed-person-parameter estimation method is applied to derive pretest item parameter estimates. The field-testing design discussed in this study is a form of seeding design commonly used in the large-scale CAT programs. Under such as seeding design, field-test (FT) items are stored in an FT item pool and a predetermined number of them are randomly chosen from the FT item pool and administered to each individual examinee. This study recommends focusing on the valid cases (VCs) that each item may end up with given a certain calibration sample size, when the FT response data are sparse, and introduces a simple strategy to identify the relationship between VCs and calibration sample size. From a practical viewpoint, when the minimum number of valid cases reaches 250, items parameters are recovered quite well across a wide range of the scale. Implications of the results are also discussed.

How to Cite

CAT Field-Test Item Calibration Sample Size: How Large is Large under the Rasch Model?. (2015). Global Journal of Human-Social Science, 15(G1), 73-79. https://socialscienceresearch.org/index.php/GJHSS/article/view/1443

References

J-C Ban, B Hanson, T Wang, Q Yi, D Harris (2001) A comparative study of on-line pretest item: Calibratoin/Scaling methods in computerized adaptive testing. 38(3), 191-212.

R Hambleton, H Swamina Than, H Rogers (1991) Fundamentals of Item Response Theory.

C Glas (2003) Quality control of online calibration in computerized assessment.

K Haynie, W Way (1995) An investigation of item calibration procedures for a computerized licensure examination.

Y Hsu, T Thompson, W.-H Chen (1998) CAT item calibration.

P Jansen, A Van Den Wollenberg, F Wierda (1988) Correcting unconditional parameter. 12(3), 297-306.

G Kingsbury (2009) Adaptive item calibration: A process for estimating item parameters within a computerized adaptive test.

J Linacre (2001) WINSTEPS Rasch measurement computer program.

H Meng, S Steinkamp (2009) A comparison study of CAT pretest item linking designs.

C Par Shall (1998) Item Development and Pretesting in a CBT Environment. 131-154.

Martha Stocking (1988) SCALE DRIFT IN ON‐LINE CALIBRATION. 1988(1).

M Stocking (1990) Specifying optimum examinees for item parameter estimation in item response theory. 55(3), 461-475.

A Van Den Wollenberg, F Wierda, P Jansen (1988) Consistency of Rasch model parameter estimation: a simulation study. 12(3), 307-313.

Wen-Chung Wang, Cheng-Te Chen (2005) Item Parameter Recovery, Standard Error Estimates, and Fit Statistics of the Winsteps Program for the Family of Rasch Models. 65(3), 376-404.

B Wright, G Douglas (1977) Best procedures for sample-free item analysis. 1, 281-295.

B Wright, M Stone (1979) Measurement, Evaluation.

M Zimowski, E Muraki, R Mislevy, R Bock (1999) BILOG-MG: Multiple group IRT analysis and test maintenance for binary items.

CAT Field-Test Item Calibration Sample Size: How Large is Large under the Rasch Model?

Published

2015-02-19

How to Cite

CAT Field-Test Item Calibration Sample Size: How Large is Large under the Rasch Model?. (2015). Global Journal of Human-Social Science, 15(G1), 73-79. https://socialscienceresearch.org/index.php/GJHSS/article/view/1443