Visualising candidate behaviour in computer-based testing: Using ClickMaps for exploring ClickStreams in undergraduate and postgraduate medical examinations
McManus, I. C.; Chis, L.; Ferro, A.; Oram, S. H.; Galloway, J.; O'Neill, V.; Myers, G.; Sturrock, A.
Show abstract
BackgroundThe rapid introduction of computer-based testing (CBT) in UK (United Kingdom) undergraduate and postgraduate medical education, mainly as a result of the COVID-19 pandemic, has generated large amounts of examination data, which we call the ClickStream. As candidates navigate through exams, read questions, view images, choose answers and sometimes change answers and return, re-read, and make further changes, the multiple actions are recorded as a series of time-stamped clicks or keystrokes. Analysing that mass of data is far from simple, and here we describe the creation of ClickMaps, which allow examiners, educationalists and candidates to visualise behaviour in examinations. MethodsAs an example of ClickMaps, we describe data from a single examination lasting three hours, with 100 best-of-five questions, which was one of two papers sat in 2021 by 508 candidates as a part of the MRCP(UK) Part 2 exam. Two ClickMaps were generated for each candidate. The Full ClickMap allows the complete three-hours of the examination to be visualised, while the Early ClickMap, shows in more detail how candidates responded during the first six minutes of presentation of each of the 100 questions in the exam. ResultsSince the primary purpose of this paper is expository, detailed descriptions and examples of ClickMaps from eleven candidates were chosen to illustrate different patterns of responding, both common and rare, and to show how straightforward are ClickMaps to read and interpret. ConclusionsThe richness of the data in ClickStreams allows a wide range of practical and theoretical questions to be asked about how candidates behave in CBTs, which are considered in detail. ClickMaps may also provide a useful method for providing detailed feedback to candidates who have taken CBTs, not only of their own behaviour but also for comparison with different strategies used by other candidates, and the possible benefits and problems of different approaches. In research terms, educationalists urgently need to understand how differences in ClickMaps relate to differences in student characteristics and overall educational performance.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Biology exams rarely use visual models to engage higher-order cognitive skills 94%
- Investigating the Role of AI Explanations in Lay Individuals’ Comprehension of Radiology Reports: A Metacognition Lense 93%
- High School Science Fair: What Students Say -- Mastery, Performance, and Self-Determination Theory 93%
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 93%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 92%
- Training Doctoral Students in Critical Thinking and Experimental Design using Problem-based Learning 91%
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 91%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 91%
Similar papers in this journal
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 92%
- Anticipatory Emotions and Academic Performance: The Role of Boredom in a Preservice Teachers' Lab Experience 90%
- Error Rates in SARS-CoV-2 Testing Examined with Bayesian Inference 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.