Back

Evaluating and Classifying Gentleness in VR-Based Surgical Simulation: A VR+fNIRS Study

Sanli, S.; Keles, H. O.

2026-01-04 neuroscience
10.64898/2026.01.02.697327 bioRxiv
Show abstract

Gentleness, defined as the ability to handle tissues delicately and minimize unnecessary force, is a key indicator of surgical proficiency. Objective and real-time assessment of gentleness in virtual reality (VR)-based training can enhance the understanding of both psychomotor and cognitive aspects of surgical skill. This study evaluates and classifies participants gentleness during VR-based surgical simulations using fNIRS-derived hemodynamic features. We trained and compared several machine learning models to assess performance. Twenty-three volunteers with no prior laparoscopic experience performed a virtual reality-based laparoscopic double-grasper task while hemodynamic activity over frontal and motor cortical areas was recorded using eighteen fNIRS channels. Alongside fNIRS, we collected subjective workload (NASA-TLX), error numbers, and a VR gentleness score. This task involves using two grasper tools simultaneously to perform the tissue like balloon manipulation in a VR environment. We extracted temporal features (slope, RMS, standard deviation) and trained machine learning models to classify performance levels based on cortical activation. Labels were binarized as low vs. high using median splits for the gentleness score. Models were evaluated with stratified 5-fold cross-validation and summarized by accuracy. Results showed stronger right-frontal HbO activity and increased left-motor HbR responses in the low-performance group, suggesting greater cognitive effort and less efficient motor strategies during VR-based laparoscopic manipulation. Across classifiers and feature sets, slope-based features consistently outperformed variability- and amplitude-based metrics. Among the tested models, HbR slope features achieved the best overall classification performance, with the highest accuracy obtained using K-Nearest Neighbor and Random Forest classifiers (accuracy {approx} 0.89, AUC up to 0.97). These findings demonstrate that fNIRS-derived hemodynamic dynamics can reliably discriminate between high and low VR performance levels, supporting their potential use in automated performance assessment and neuroadaptive feedback frameworks for VR-based surgical training.

Published in Sensors (predicted rank #8) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.