Exploring Educational Efficiency:
A Comprehensive Statistical Analysis of Class-Level Metrics, Question Types, and Participant
Performance
Abstract
This study offers a rigorous, multifaceted investigation of participant performance and class
efficiency in a university-level educational setting. Drawing on data from three faculties and
involving 119 participants, the research integrates one-way ANOVA, Tukey’s HSD post-hoc
tests, correlation analysis, regression modelling, clustering, and question-type examinations
to unpack the complexities underlying learning outcomes. Results from the ANOVA confirm
statistically significant performance disparities across multiple classes, while subsequent
post-hoc comparisons pinpoint precise class pairs that exhibit marked accuracy gaps.
Correlation and regression analyses reveal that merely spending more time on tasks—often
presumed to improve achievement—does not consistently lead to higher accuracy; instead,
participant knowledge (as captured by score) and specific class contexts wield a more
pronounced influence. Additional insights emerge from the question-type evaluations,
indicating that intensive assessment formats (e.g., Reorder or Open-Ended) can yield higher
accuracy but demand greater time investment, whereas simpler tasks provide rapid
responses but risk lower correctness rates. K-Means clustering, visualised via Principal
Component Analysis, further demonstrates the heterogeneity of learner profiles: some
participants achieve high accuracy while expending relatively little time, whereas others
excel only through protracted effort. Grounded in established educational theories—
including Bloom’s taxonomy and socio-cultural perspectives—these findings illuminate the
interplay between individual learner characteristics, instructional design, and broader
contextual factors. By synthesising diverse analytic approaches within a single framework,
the study presents actionable strategies for refining teaching methods, optimising
assessment formats, and enhancing educational outcomes across heterogeneous learning
environments.
Introduction
Educational research has long emphasised the importance of aligning teaching methods,
assessment strategies, and learner characteristics to foster meaningful academic progress
(Biggs, 2011; Anderson & Krathwohl, 2001). Early cognitive theories, exemplified by Bloom’s
taxonomy, posit that knowledge acquisition progresses from basic recall to higher order
thinking skills, underscoring the need for pedagogical approaches and assessment items
tailored to learners’ developmental stages (Krathwohl, 2002). Socio-cultural perspectives,
such as Vygotsky’s (1978) concept of the Zone of Proximal Development, further underscore
the context-dependent nature of learning, suggesting that social interaction and carefully
structured scaffolding play pivotal roles in student performance and motivation. Over the
last decade, the emergence of learning analytics has facilitated a more data-driven
understanding of how students engage with course materials and assessments (Siemens &
Long, 2011). This trend has been reinforced by advances in educational data mining,
enabling researchers to analyse large-scale participant datasets and identify patterns in
performance, time management, and engagement (Baker & Yacef, 2009). Yet studies
investigating the relationship between time on task and accuracy have yielded mixed
findings, with some revealing modest positive correlations (Kim et al., 2018) and others
reporting null or even negative associations (Tempelaar et al., 2015). These discrepant
outcomes reinforce the complexity of classroom dynamics and call attention to additional
variables—such as question difficulty, learner motivation, and instructional methods—that
may exert substantial influence on educational outcomes (Gasevic et al., 2015). In parallel,
earlier scholarship highlights how question design can shape both the depth of student
understanding and overall test performance (Mayer, 2002). Multiple-choice questions, for
instance, efficiently gauge factual knowledge but may underestimate higher-order reasoning
(Rushton, 2005). Conversely, open-ended, or project-based tasks can foster deeper learning
but demand more time and nuanced feedback (Reeves, 2006). As classrooms grow
increasingly diverse—in both learner backgrounds and technological aptitudes—
researchers have argued for multifaceted analytical strategies that integrate traditional
statistics (e.g., ANOVA, correlation) with more powerful machine learning or clustering
techniques, thereby capturing the full range of variability in student performance (ZawackiRichter et al., 2019). Building on these theoretical foundations, the present study adopts a
unified framework combining one-way ANOVA, Tukey’s HSD post-hoc tests, correlation
analysis, regression modelling, clustering, and question-type examinations to investigate
participant performance and class efficiency. Situated within established educational
theories and leveraging advanced data analytics, this work aims to elucidate how multiple
factors—ranging from time management to assessment design—interact to shape accuracy
and overall learning outcomes. The subsequent Methods section details the data collection
process, variables of interest, and analytic procedures, paving the way for an in-depth
examination of class-level differences, time–accuracy relationships, and the myriad ways
participants engage with various assessment formats. Nonetheless, while the current data
offer a thorough analysis of performance across classes, question types, and learner profiles,
additional factors—such as participant demographics, motivational drivers, and curricular
structures—could further enrich these interpretations.
Methods
All data analysed in this research were obtained from participants’ interactions with an
online assessment platform, which tracked multiple performance indicators in real time.
Five key metrics were collected for each participant:
1. Accuracy (percentage of correct responses), which serves as a straightforward measure of
correctness.
2. Total Time Taken (recorded in seconds), reflecting the duration each participant spent
responding to items; this metric can serve as a proxy for cognitive effort or pacing strategies.
3. Score, representing cumulative points accrued based on correctness and speed; depending
on the platform’s rules, partial credit or additional points for rapid answers may be involved
(e.g., +10% for answering in under 10 seconds, partial credit for partially correct reorder
questions, etc.).
4. Class Name, identifying the instructional cohort to which each participant belongs.
5. Question Type, denoting item formats (e.g., Multiple Choice, Fill-in-the-Blank, Reorder)
that can vary in complexity and cognitive demands.
Collectively, these metrics enabled a granular examination of how participant-level (e.g.,
Accuracy, Total Time Taken, Score) and question-level (e.g., Question Type) factors converge
to shape learning outcomes. The inclusion of both performance and process variables aligns
with recommendations in the educational data mining literature (Baker & Yacef, 2009).
Specifically, while Accuracy captures a broad indicator of success on given tasks, Total Time
Taken reveals differences in pacing and effort, and Score accounts for speed-based
incentives or partial-credit systems. Similarly, Question Type sheds light on item complexity,
which can influence both the depth of cognitive engagement and time requirements. By
integrating these metrics, the study aimed to isolate the effects of class membership,
question design, and individual pacing on overall achievement, thereby providing a multidimensional view of educational efficiency. This approach underscores the importance of
capturing not only whether students succeed, but also how they interact with assessment
items, a distinction crucial for refining pedagogical strategies and improving learner
outcomes.
Data Collection and Preparation
Data for this study were gathered during the winter semester of the 2024/2025 academic
year at three distinct faculties: the Faculty of Metallurgy, the Faculty of Science and
Mathematics, and the Faculty of Philology. This targeted selection was intended to ensure
broad disciplinary representation, capturing students with varied academic interests and
subject specialisations. Building on the idea that learning dynamics often differ substantially
across scientific, technical, and humanities-oriented disciplines (Biggs, 2011), these three
faculties provided a more comprehensive vantage point for understanding how instructional
contexts and learner characteristics intersect to influence educational performance. The
cohorts under investigation included SUK (Oral Communication Strategies), SEJ 9
(Contemporary English Language), and PRE (Translation Studies), representing first- and
second-year master’s students at the Faculty of Philology; PMF 1 (General English 1) and
PMF 3 (English for Specific Purposes), representing first- and second-year students at the
Faculty of Science and Mathematics; and MTF 1 (General English), composed of first-year
students at the Faculty of Metallurgy. Altogether, 119 participants were included, ranging in
age from 18 to 39, with a mean age of 19.7 years. Informal logs indicated approximate cohort
sizes of SUK (n=20), SEJ 9 (n=17), PRE (n=15), PMF 1 (n=25), PMF 3 (n=20), and MTF 1
(n=22)—though exact numbers may vary slightly due to absences or incomplete data. This
wide age span encompassed both traditional and non-traditional students, diversifying the
sample and thereby enhancing the external validity of the findings.
Prior to formal analyses, the dataset underwent systematic cleaning to ensure accuracy and
completeness. Entries were examined for potential anomalies (e.g., negative time values,
missing class labels), and outliers in total time were flagged when they appeared implausibly
high or low—suggesting possible data-capture errors. Although no strict threshold was
predefined, extreme values were scrutinised individually and, if necessary, removed, or
capped (e.g., at the 99th percentile) to limit distortions in subsequent statistical modelling.
Once the data were verified as reliable, descriptive statistics were calculated for Accuracy,
Total Time Taken, and Score. Visual diagnostics such as histograms and boxplots then
assisted in identifying whether these variables followed near-normal or skewed
distributions, which is especially relevant for validating assumptions behind regression
models and ANOVA. This preparation ensured that the dataset was robust enough to support
the analytical techniques employed, bolstering confidence in the validity of the ensuing
results.
One-Way ANOVA and Tukey’s HSD
A one-way ANOVA was first employed to assess whether mean accuracy differed
significantly across classes. Specifically, we tested the hypothesis that H₀: all class means are
equal vs. H₁: at least one class mean is different, using an alpha level of .05. The dataset
included six instructional cohorts, thus giving rise to five numerator degrees of freedom. The
overall ANOVA produced F(5, 113) = 53.28, p < .0001, partial η² = 0.29 (Biggs, 2011). This
partial η² suggests a large effect according to conventional benchmarks, implying that
around 29% of the variance in Accuracy is attributable to between-class differences. While
a significant F-ratio indicates that at least one class diverges statistically from the overall
mean, it does not pinpoint which groups differ most. Consequently, we employed a Tukey’s
HSD post-hoc test, controlling the familywise error rate at α = 0.05 for the 15 pairwise
comparisons among six classes (Anderson & Krathwohl, 2001). This two-step procedure—
ANOVA followed by Tukey’s HSD—helps limit Type I error inflation. Widely accepted in both
educational and psychological research, it pinpoints which class-level discrepancies are
meaningful, offering clearer insights into how different instructional cohorts compare in
terms of accuracy.
Correlation Analysis
Pearson correlation coefficients were calculated within each class to explore potential
relationships between time and accuracy. This measure of linear association clarifies
whether increased time usage generally co-occurs with higher accuracy (positive
correlation) or lower accuracy (negative correlation). However, as the literature in learning
sciences indicates, time on task does not invariably translate into improved performance
(Tempelaar et al., 2015). Some learners employ extra time effectively—e.g., through review
and reflection—whereas others may drift off-task or use time inefficiently. By conducting
correlations at the class level, the study acknowledges that pedagogical strategies, learner
demographics, and other contextual factors could differentially shape the time–accuracy
relationship (Gasevic et al., 2015). In addition to the correlation coefficient r, we also
computed p-values for each correlation and a 95% confidence interval. Classes PMF 1 and
SEJ 9 showed notable negative correlations (r ≈ –0.19 and –0.20, respectively), both
significant at p < .01, suggesting that in these cohorts, spending more time might even
correlate with slightly lower accuracy—though effect sizes remained modest (r² < 0.05).
Such findings underscore the complexity of relating time usage to performance, hinting that
certain classes may foster time-efficient strategies or that struggling students simply spend
longer without improving results.
Regression Models
Building on the correlation findings, two regression models were formulated to account for
variability in Accuracy:
1. Simple Linear Regression
\[
\text{Accuracy} \sim \text{Total Time Taken}
\]
This baseline model assessed the explanatory power of time alone, testing the assumption
that “more time on task equals better outcomes.” A negligible or negative slope would
contradict the notion that prolonged engagement inevitably boosts performance. The
resulting slope was –0.0025 (p = 0.006), with an R² = 0.003, reflecting minimal explanatory
power.
2. Multiple Regression
\[
\text{Accuracy} \sim \text{Score} + \text{Total Time Taken} + \text{Class Name}
\]
This second, more comprehensive model added both knowledge-related (Score) and
contextual (Class Name) variables. Since performance rarely hinges on time alone (Gasevic
et al., 2015), the inclusion of Score (often reflecting speed, correctness, and partial credits)
and class membership provided a richer account of observed variations. The final model
achieved R² = 0.268 (adjusted R² = 0.266), partial F-tests indicated Score (p < .001) as the
strongest predictor, and certain class-level variables (Class Name_PMF 1, Class Name_SUK)
also being significant at p < .001. While no severe multicollinearity emerged (VIF < 2 for all
predictors), residual diagnostics suggested slight curvature (see Figure 9 below) and mild
heteroskedasticity, implying that polynomial or interaction terms (e.g., Time × Score) might
further improve fit. Nonetheless, the model’s improvement over the simple regression
underscores the importance of knowledge level and class context relative to time alone.
Question-Type Analysis
The question-type analysis investigated how item format influenced both accuracy and time.
It stems from the principle that different question formats demand varying levels of
cognitive processing (Mayer, 2002). For instance, Reorder items could require participants
to grasp sequences or logical structures, potentially resulting in greater time usage yet
deeper engagement and, consequently, improved accuracy. Conversely, tasks like Multiple
Choice may rapidly gauge factual knowledge but lack the capacity to measure advanced
reasoning (Rushton, 2005). While existing data revealed general trends, further item-level
details—e.g., difficulty indices, Bloom’s taxonomy categorisation—could refine these
observations. Such specificity might differentiate tasks that truly foster deeper
comprehension from those that simply take longer due to logistical complexities.
Additionally, partial-credit rules for certain formats (e.g., awarding half points for partially
correct reorder tasks) may influence how Score relates to Accuracy, particularly if
participants earn partial credit without fully completing an item. Understanding how these
scoring rules operate is crucial: if Reorder items yield more partial credit opportunities, a
student might accumulate Score points despite not fully mastering the content, potentially
affecting correlations between time, accuracy, and Score.
K-Means Clustering and PCA
Finally, K-Means clustering was employed to group participants based on accuracy, score,
and total time. We used the elbow method on the within-cluster sum of squares to determine
that k=3 offered a reasonable balance between model simplicity and explanatory power
(Baker & Yacef, 2009). Clustering can uncover latent performance profiles—for instance,
learners who excel quickly versus those who require more extensive engagement. Such
identification resonates with emerging calls for adaptive education, wherein recognising
diverse learner archetypes can inform tailored pedagogical interventions.
[Figure 8: Participant Clusters (PCA-Reduced) here]
Principal Component Analysis (PCA) was used to project the clustering results onto two
principal components, explaining approximately 65% of the variance in the threedimensional space (Accuracy, Score, Time). This facilitated a more intuitive visual separation
among groups (Siemens & Long, 2011). Each cluster’s size ranged from Cluster 0 (n=30) to
Cluster 1 (n=40) to Cluster 2 (n=49), indicating that no single cluster dominated the sample.
By integrating ANOVA, post-hoc testing, correlation, regression, question-type comparisons,
and clustering, the study addresses the multifaceted nature of educational efficiency, from
class-level dynamics to individual time-management strategies. In doing so, it aims to weave
together each analytic method into a coherent narrative, presented in the next section,
illuminating the myriad factors that shape performance.
Results
The analyses conducted in this study illuminate distinct dimensions of participant
performance, including class-level disparities, the relationship between time and accuracy,
variations in question-type outcomes, and the formation of unique participant clusters. The
following sections detail each of these findings, accompanied by references to the
corresponding figures. Where needed, numeric results are contextualised with interpretive
commentary, linking them to broader educational theories and practical implications.
Class-Level Comparisons and Tukey’s HSD
A one-way ANOVA tested whether mean accuracy differed across classes, yielding F(5, 113)
= 53.28, p < 0.0001, partial η² = 0.29. These differences become visually evident in [Figure 1:
Tukey HSD: Class Accuracy Comparisons], which displays each class’s mean accuracy and
confidence intervals. Additional detail on specific pairwise contrasts is provided by [Figure
2: Mean Differences Between Classes (Tukey HSD)] and [Figure 3: Significant Class
Comparisons (Tukey HSD)], clarifying which pairs of classes diverge most significantly.
Among the largest disparities were MTF 1 vs. PMF 1 (difference = 13.69, p < 0.001), PMF 1
vs. SEJ 9 (–14.95, p < 0.001), and SEJ 9 vs. SUK (11.82, p = 0.048).
[Figure 13: Mean Accuracy by Class] further places PMF 1 at approximately 79% accuracy,
SEJ 9 around 60%, SUK near 73%, MTF 1 at 65%, PMF 3 at 70%, and PRE at 68%. These
results underscore PMF 1’s notably high performance. The pronounced differences suggest
that certain structural or pedagogical factors unique to PMF 1 may be driving its superior
outcomes, while SEJ 9’s relative underperformance calls for closer examination of its
instructional methods or learning environment. From a theoretical standpoint, these
disparities may reflect differing applications of instructional design principles (Ertmer &
Newby, 1993) or varying degrees of alignment with deeper learning objectives (Rosenshine,
2012). For instance, classes demonstrating higher accuracy may employ collaborative
activities, well-structured feedback loops, or scaffolds that facilitate mastery. Conversely,
classes with lower outcomes might experience disjointed pacing, insufficient conceptual
reinforcement, or other instructional gaps.
Time Usage Among Classes
Time allocation across classes was similarly revealing. In [Figure 11: Class-Level
Performance: Accuracy vs. Time], accuracy and total time are combined into a stacked
visualisation, showing, for example, that PMF 3 exceeds 400 seconds on average yet remains
near 70% accuracy, while PMF 1 expends approximately 320 seconds but attains higher
accuracy. A separate focus on time alone appears in [Figure 12: Mean Total Time by Class],
where PMF 3 surpasses 480 seconds on average and SEJ 9 nears 450 seconds, whereas PMF
1 remains closer to 320 seconds. These findings demonstrate that greater time investment
does not necessarily translate into superior results, suggesting that some classes optimise
both pacing and precision more effectively than others. This discrepancy parallels the
broader debate on the relationship between time on task and learning quality (Biggs, 2011).
While sufficient time is crucial for tasks requiring complex reasoning, time alone cannot
guarantee successful outcomes if students lack strategic learning behaviours or if
instructional design fails to guide them effectively. In other words, the data highlight a “time–
quality” paradox: more time is necessary for certain tasks, yet it is not always exploited in a
manner that boosts accuracy. Additional contextual details—such as how instructors
scaffold tasks or how feedback is integrated—might offer further clues. If such data are
available (e.g., logs of how participants used their time or teacher reflections on pacing), they
could elucidate why some classes thrive with less total time while others appear less
efficient.
Correlation and Regression Insights
To assess whether prolonged engagement leads to higher accuracy, a scatterplot of total time
versus accuracy for individual participants was constructed. [Figure 4: Accuracy vs. Total
Time Taken] shows that increases in time do not reliably predict improvements in
correctness; indeed, PMF 1 and SEJ 9 both display negative correlations (–0.19 and –0.20,
respectively), despite SEJ 9 allocating more time overall. A simple linear regression of
Accuracy on Total Time confirmed this pattern, producing an R-squared of 0.003 and a
slightly negative slope (–0.0025, p = 0.006), as illustrated in [Figure 5: Regression: Total
Time vs. Accuracy].
Introducing additional predictors improved model performance. In the multiple regression
Accuracy ~ Score + Total Time Taken + Class Name, R-squared rose to 0.268 (adjusted
0.266), with partial eta-squared for Score indicating a strong effect (p < .001). Certain classlevel variables also showed significance (Class Name_PMF 1 = 9.27, p < 0.001; Class
Name_SUK = 16.05, p < 0.001). Nevertheless, [Figure 9: Residuals vs. Fitted Values] suggests
slight curvature in the residuals, signalling that further nonlinearities or interactions might
remain unaccounted for. Additionally, [Figure 10: Actual vs. Predicted Accuracy] illustrates
variability around the diagonal of perfect prediction, especially among higher-accuracy
participants. Collectively, these findings challenge the assumption that “time on task”
equates directly to improved performance (Tempelaar et al., 2015). Instead, knowledge
levels (Score) and class contexts seem to overshadow the predictive value of time alone. One
might hypothesise that classes employing effective metacognitive training or scaffolded
assignments can achieve strong results in less time, whereas other groups spend longer but
with more disorganised or less productive study behaviours. Future studies incorporating
qualitative insights—such as interviews or in-class observations—could clarify how
participants invest their time, distinguishing between passive re-reading and active
problem-solving strategies.
Question-Type Analysis
Assessment format emerged as another significant factor. As shown in [Figure 6: Mean
Accuracy by Question Type], items such as Reorder, Open Ended, and View Player Data
regularly achieve over 80–90% correctness, whereas Check Box items linger near 40%. This
difference in performance can be understood in light of [Figure 7: Average Time Taken by
Question Type], revealing that more complex tasks (e.g., Comprehension_Video at ~50
seconds) consume considerably more time, while simpler question types (e.g., Check Box)
take around 20 seconds. Educators thus face a clear trade-off: in-depth question formats
tend to promote more accurate responses but also demand additional time. These findings
resonate with research on cognitive load and student engagement (Mayer, 2002). While
more demanding questions can foster deeper processing, there is a risk of overwhelming
learners who lack the requisite background knowledge or time-management skills.
Meanwhile, rapid-response tasks, although quick and efficient, may fail to capture students’
higher-order thinking capacities. Striking a balance between these question types becomes
essential for instructors aiming to evaluate a broad spectrum of cognitive skills. In practice,
an instructor might design a mixed-format assessment combining lower-order, rapid checks
for basic recall with a smaller number of high-level tasks that probe synthesis or evaluation
(Anderson & Krathwohl, 2001). If further data are available on item difficulty, question-level
hints, or Bloom’s taxonomy categorisations, future work could differentiate “hard Reorder”
from “easier Reorder” items, clarifying whether time-intensive question types consistently
yield deeper learning or simply longer tasks without additional benefits.
Participant Clusters
Finally, K-Means clustering offered a deeper look at patterns among individuals, segmenting
participants based on Accuracy, Total Time, and Score. [Figure 8: Participant Clusters (PCAReduced)] presents these clusters in two-dimensional PCA space, revealing three distinct
groups. One cluster exhibits moderate accuracy (~58.88%), moderate scores (~9404.84),
and lower time usage (~315.88 seconds). A second cluster achieves high accuracy
(~84.20%), high scores (~21405.89), and moderate time (~351.84 seconds). The third
cluster also attains high accuracy (~81.66%) but does so through a much longer average
time expenditure (~3610.94 seconds) and a comparatively lower score (~6019.69). Notably,
the presence of speed-based bonuses in the scoring formula appears to penalise slower
respondents, even if they achieve high accuracy. Each cluster included between 60–90
participants, indicating a relatively balanced partition. These clusters highlight the
heterogeneity in how learners approach tasks. Cluster 0’s moderate accuracy and moderate
score combined with lower time usage might reflect learners who possess basic competence
yet do not deeply engage, whereas Cluster 1’s high scores, high accuracy, and moderate time
may indicate well-prepared individuals who manage pacing effectively. Cluster 2’s high
accuracy but low score, coupled with extended time usage, could point to a subset of learners
who attempt fewer items or proceed slowly but meticulously—thus boosting correctness
but failing to accumulate speed-based points. From an instructional design viewpoint, these
findings support differentiated instruction (Ertmer & Newby, 1993). Rapid-accuracy
learners might benefit from challenge extensions or advanced problems to stay motivated,
whereas slower but accurate learners could require guidance in time management or
strategic question approaches to avoid being penalised. Importantly, no single cluster is
“better” than another; each represents a distinct performance profile with its own
advantages and challenges. By recognising this diversity, educators can deploy interventions
targeting specific clusters, personalising the learning experience and potentially improving
overall outcomes. Overall, the results illuminate how class context, time management,
question design, and distinct learner profiles intersect to shape performance outcomes. The
finding that certain classes surpass others in both accuracy and efficiency, that question
types yield contrasting balances of speed and correctness, and that participant clusters
follow multiple effective pathways collectively underscore the importance of employing
multiple analytical tools to capture the full breadth of educational efficiency.
Discussion
The results of this study highlight the multilayered and interdependent nature of educational
performance, demonstrating that class-level pedagogical strategies, learner characteristics,
and question design all substantially shape outcomes. Although PMF 1’s consistently high
accuracy suggests the value of replicating its teaching practices or curricular structures,
these benefits may be context-dependent or influenced by demographics and resource
allocation. The marked differences in mean accuracy among classes, especially in the
contrast between PMF 1 and SEJ 9, indicate that underperforming groups require a targeted
analysis of local learning conditions—ranging from instructional pacing and feedback
mechanisms to broader classroom culture.
Interpreting Class-Level Disparities
One of the most salient findings is the notable accuracy gap across classes, aligning with the
concept of constructive alignment (Biggs, 2011). PMF 1’s strong performance could be
attributed to well-aligned learning objectives, consistent scaffolding practices, or effective
feedback loops that guide students toward mastery. In contrast, SEJ 9’s comparatively
weaker outcomes may stem from misalignment between instructional methods and
assessment tasks, insufficient instructor–student interaction, or contextual variables such as
student motivation or socio-economic status. Future research might illuminate these
possibilities through classroom observations or qualitative data collection, helping to
differentiate how each factor influences learners. From a systemic perspective, class-level
disparities may also reflect broader inequities in resource distribution (Terenzini &
Pascarella, 1994). For instance, institutions offering more advanced technological support
or smaller student–teacher ratios often see improved academic results. Thus, future analyses
integrating class size, teacher qualifications, and resource availability would be beneficial for
clarifying the roots of these performance gaps. By scrutinising such contextual data,
stakeholders can design more effective interventions that address both micro-level
classroom practices and macro-level institutional policies.
The Time–Accuracy Paradox
The observation that time usage alone is often weakly or even negatively correlated with
accuracy corroborates earlier findings that simply extending test-taking periods does not
invariably enhance performance (Tempelaar et al., 2015). However, this negative association
indicates a more nuanced reality: some learners may use additional time ineffectively—
perhaps re-reading questions without employing strategic review—whereas others
leverage brief yet focused efforts to achieve strong results. The regression analyses reinforce
that knowledge levels (captured by Score) and class contexts better explain outcome
variations than time alone. From a self-regulated learning lens (Ertmer & Newby, 1993), time
can be beneficial if learners adopt reflective strategies, break down complex tasks, or actively
engage with feedback. However, lacking such strategies, extended time may devolve into
unproductive procrastination or guesswork. These findings have considerable implications
for how institutions and instructors set deadlines, structure assessments, and integrate
formative feedback. Prioritising training in time management or metacognitive skills might
prove more effective than merely offering students additional minutes on a quiz or exam.
Question-Type Complexities
The question-type analysis further clarifies how assessment format influences both accuracy
and time demands. Higher-accuracy item types (e.g., Reorder, Open-Ended) often promote
deeper cognitive engagement, consistent with theories emphasising active mental
processing (Mayer, 2002). Yet, these formats also require more time, creating a trade-off that
educators must navigate when designing assessments. A balanced approach might involve
combining cognitively demanding tasks that reinforce higher-order thinking (analysis,
synthesis, evaluation) with more rapid-response items that check basic comprehension. This
aligns with Bloom’s taxonomy (Anderson & Krathwohl, 2001), which underscores that tasks
testing advanced skills are inherently more time intensive. While such items can foster
deeper mastery and real-world problem-solving skills, they may also strain learners if not
carefully scaffolded. Differentiating formative tasks (aimed at diagnostic or developmental
feedback) from summative tasks (aimed at certification of competence) can help educators
mitigate the risk of overburdening students while still promoting meaningful learning gains.
Heterogeneity of Learners and Clustering Insights
The K-Means clustering analysis reveals that participants vary widely in how they balance
time and accuracy. Some achieve high accuracy rapidly, others succeed through protracted
engagement, and still others exhibit moderate achievement with moderate time usage. Such
heterogeneity underscores the need for adaptive instructional strategies, particularly in
large and diverse classrooms. Time-management workshops, strategic-question
approaches, or additional enrichment activities could benefit different learner clusters,
acknowledging that no single pathway to success is universal. These clustering insights
resonate with personalised learning frameworks (Wang & Wu, 2020), suggesting that
interventions might be fine-tuned according to learner “types.” A well-prepared yet timeefficient group might need more challenging material to maintain motivation, whereas
methodical but slower learners may benefit from explicit time-management guidance or
additional scaffolding. If further data on learner motivation, socio-economic status, or prior
academic achievement become available, future clustering analysis could yield even more
precise subgroups, thereby paving the way for highly targeted educational interventions.
Limitations and Future Directions
Despite this study’s rigorous analytical approach, several limitations warrant mention. First,
the assumption that total time accurately reflects genuine engagement may not always hold;
students might become idle, multitask, or face technical interruptions, leading to
discrepancies in the recorded time. Second, the dataset does not incorporate crucial
variables such as learner motivation, metacognitive strategy use, or demographic factors
(e.g., socio-economic status). Each of these could significantly shape performance and might
explain inter-class differences beyond those captured in the current models. Third, while the
question-type analysis offers meaningful insights, incorporating an item-level difficulty
metric or mapping each item to a taxonomy-based classification could enhance interpretive
granularity. In addressing these gaps, future research might collect additional “process data”
(e.g., clickstream logs, dwell times on individual items, user-initiated hints) to form a more
comprehensive view of participant engagement. Integrating measures of intrinsic or
extrinsic motivation, such as brief survey instruments, could reveal the degree to which time
on task correlates with deeper engagement or procrastination-driven behaviours. Multilevel modelling could further disentangle the proportion of variance attributable to
individual learners versus instructional environments, thereby clarifying how teaching
practices intersect with learner traits. Finally, qualitative examinations of high-performing
classes, notably PMF 1, could yield best-practice insights—covering everything from teacher
expertise to the pacing of lessons or peer collaboration strategies. Adapting these best
practices for lower-performing cohorts, such as SEJ 9, may serve as an impactful step toward
elevating overall educational outcomes. Through this multi-pronged lens—linking time
usage, question design, learner characteristics, and class contexts—the present study
illuminates the complexity of fostering effective learning. Addressing these limitations and
pursuing future lines of inquiry will contribute to a more holistic, evidence-based
understanding of educational efficiency and performance.
Conclusion
The findings from this multifaceted analysis affirm that educational outcomes result from a
complex interplay of class-level strategies, nuanced assessment designs, and individual
learner pathways. One-way ANOVA (F(5, 113) = 53.28, p < .0001, partial η² = 0.29) and
Tukey’s HSD comparisons demonstrate that certain classes, notably PMF 1, achieve
consistently higher accuracy while investing less time, suggesting that replicating successful
pedagogical or structural elements could enhance performance across contexts. In contrast,
SEJ 9 spends substantially more time without proportional gains in correctness,
underscoring that merely allocating extra minutes does not inherently boost achievement.
Correlation and regression analyses reinforce this perspective by revealing that total time
alone exerts minimal explanatory power, whereas score (indicative of knowledge level) and
class-related factors more robustly predict accuracy. Nevertheless, residual diagnostics and
the presence of slight nonlinearities suggest that no single variable or linear model can fully
capture the multi-dimensional processes of learning—hinting at underexplored
motivational, cognitive, and contextual influences. A deeper understanding of these
complexities emerges in the question-type analysis and K-Means clustering. More timeintensive question formats can indeed improve accuracy but may impose cognitive and
pacing demands unsuited to all learners. The clustering results further show that high
accuracy can emerge from multiple strategic approaches—some learners excel with swift
precision, while others navigate tasks methodically to reach comparably strong outcomes.
Collectively, these findings underscore the need for evidence-based, context-sensitive
refinements in instructional methods, from pinpointed interventions and adaptive time-
management guidance to thoughtful balancing between speed-oriented tasks and
conceptually demanding question formats. By embracing the intricacies unveiled through
residual diagnostics, scatterplots, question-type trade-offs, and diverse clustering profiles,
educators and institutions stand to create learning environments and assessments that more
effectively address students’ varied trajectories. Ultimately, this integrative, data-informed
approach promises to enhance both academic performance and transferable learning across
a wide spectrum of educational contexts.
Practical Implications
Beyond their theoretical significance, these findings present concrete pathways for
improving educational practice. The marked class-level differences and diverse learner
profiles highlight the importance of differentiated instruction, where materials and tasks are
tailored to learners who are either quick-and-accurate or slow-and-meticulous, ensuring
optimal support for each profile. The question-type analysis underscores a speed–depth
trade-off: while multiple-choice or check-box items can efficiently gauge baseline
knowledge, higher-order tasks like open-ended or reorder questions may foster deeper
thinking but demand more time. A balanced assessment strategy, blending both rapidresponse and conceptually rich items, can capture a broader range of cognitive skills without
overwhelming learners. Since score and class context emerged as stronger predictors than
time, instructors should prioritise content mastery through scaffolding, timely feedback, and
extensive practice opportunities rather than merely extending test durations or allowing
unlimited attempts. The negative or negligible correlation between time and accuracy
indicates that some students may need explicit guidance on time management, strategic
question tackling, and meta-cognitive reflection. Both high- and low-performing groups
could benefit from workshops or structured practice sessions in these areas. Moreover,
classes that chronically underperform (such as SEJ 9 in this dataset) may require focused
institutional support, ranging from smaller class sizes and additional teaching assistants to
better resource allocation. Emulating successful methods from PMF 1—where accuracy is
robust and time usage moderate—could uplift overall performance through a more
consistent alignment of objectives, instruction, and assessment. Such targeted interventions,
when systematically evaluated and refined, hold the potential to enhance instructional
quality and promote more equitable learning outcomes across diverse educational settings.
References
Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and
assessing: A revision of Bloom's taxonomy of educational objectives. Longman.
Baker, R. S., & Yacef, K. (2009). The state of educational data mining in 2009: A review and
future visions. Journal of Educational Data Mining, 1(1), 3–17.
Biggs, J. (2011). Teaching for quality learning at university: What the student does. McGrawHill Education (UK).
Ertmer, P. A., & Newby, T. J. (1993). Behaviorism, cognitivism, constructivism: Comparing
critical features from an instructional design perspective. Performance Improvement
Quarterly, 6(4), 50–72.
Gasevic, D., Dawson, S., & Siemens, G. (2015). Let’s not forget: Learning analytics are about
learning. TechTrends, 59(1), 64–71.
Kim, J., Park, H., & Cozart, J. (2018). Affective and motivational factors of learning in online
mathematics courses. British Journal of Educational Technology, 49(2), 369–382.
Krathwohl, D. R. (2002). A revision of Bloom’s taxonomy: An overview. Theory into Practice,
41(4), 212–218.
Mayer, R. E. (2002). Multimedia learning. In Psychology of Learning and Motivation (Vol. 41,
pp. 85–139). Academic Press.
Reeves, T. C. (2006). How do you know they are learning? The importance of alignment in
higher education. International Journal of Learning Technology, 2(4), 294–309.
Rosenshine, B. (2012). Principles of instruction: Research-based strategies that all teachers
should know. American Educator, 36(1), 12–39.
Rushton, A. (2005). Formative and summative assessment: Perceptions and realities. Nurse
Education Today, 25(4), 357–364.
Siemens, G. (2013). Learning analytics: The emergence of a discipline. American Behavioral
Scientist, 57(10), 1380–1400.
Siemens, G., & Long, P. (2011). Penetrating the fog: Analytics in learning and education.
EDUCAUSE Review, 46(5), 30–32.
Tempelaar, D. T., Rienties, B., & Giesbers, B. (2015). In search for the most informative data
for feedback generation. Computers in Human Behavior, 47, 157–167.
Terenzini, P. T., & Pascarella, E. T. (1994). Living with myths: Undergraduate education in
America. Change: The Magazine of Higher Learning, 26(1), 28–32.
Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes.
Harvard University Press.
Wang, Y., & Wu, M. (2020). Adaptive learning systems and their applications in education.
Computers in Human Behavior, 113, 106526.
Zawacki-Richter, O., Marin, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of
research on artificial intelligence applications in higher education. International Journal of
Educational Technology in Higher Education, 16(1), 39.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )