Machine Scoring

Machine Scoring

Not All AI Scoring Is Created Equal

As Artificial intelligence (AI) becomes more common in education, schools need more than automated scoring. They need evidence they can trust.

ACTFL® and Language Testing International® (LTI) work closely to ensure our machine scoring solutions are informed by ethic standards and research, trained on thousands of human-scored responses, aligned to ACTFL Proficiency Guidelines, and continuously monitored for accuracy and fairness.

Built on standards. Validated by human experts. Backed by evidence.

Explore the Research

WHY EDUCATORS SHOULD LOOK BEYOND THE AI LABEL

Artificial intelligence is becoming increasingly common in assessment, but technology alone does not guarantee accuracy.

Reliable machine scoring requires:

  • Extensive language-specific training data
  • Alignment to recognized proficiency standards
  • Validation against expert human raters
  • Ongoing monitoring and quality assurance
  • Ethical oversight and transparency

The quality of the evidence, and not the appeal of the technology, should guide assessment decisions.

WHY SCHOOLS AND PROGRAMS CAN TRUST ACTFL & LTI MACHINE SCORING

Developed for Accuracy. Designed for Trust.

Nearly a decade of research and development

Built over years of data collection, validation, testing, and refinement before operational deployment.

Trained on ratings by ACTFL-certified raters

The scoring model learns from thousands of responses previously evaluated by ACTFL-certified human raters.

Aligned to ACTFL Standards

Scoring is grounded in the ACTFL Proficiency Guidelines and ACTFL Performance Descriptors.

Validated before implementation

Machine-generated scores were compared against human ratings to verify scoring accuracy and consistency before going live.

Continuously monitored and improved

Performance is regularly reviewed, calibrated, and evaluated to maintain quality and accuracy over time.

Supported by independent expert oversight

An independent Expert Review Committee provides guidance on performance, monitoring, ethics, and appropriate use.

MACHINE SCORING IS NOT GENERATIVE AI

Generative AI creates content. Machine scoring evaluates language performance against established scoring criteria.

ACTFL does not employ AI for content generation on any ACTFL tests.

ACTFL and LTI's machine scoring technology is designed to mirror validated human judgment using recognized proficiency standards, not to generate responses or make subjective decisions.

BUILT ON HUMAN EXPERTISE

Machine scoring starts with human scoring.

The system is trained using authentic student responses previously scored by ACTFL-certified human raters. Through this process, the model learns the language features and performance characteristics associated with specific proficiency levels.

The goal is simple: provide scoring that is consistent, scalable, and closely aligned with expert human judgment.

FAIRNESS AND ETHICS ARE BUILT INTO THE PROCESS

Responsible machine scoring requires more than technical performance.

ACTFL’s and LTI's approach includes representative training data, ongoing validation, independent oversight, and adherence to established language assessment ethics principles. These safeguards help support fairness, transparency, and assessment integrity.

RESEARCH & EVIDENCE

ACTF’s and LTI's machine scoring approach is grounded in years of research, validation, and continuous oversight. Explore the studies, validation research, conference presentations, and technical papers documenting the evidence behind responsible machine scoring in world language assessment.

View Research Library

FREQUENTLY ASKED QUESTIONS ABOUT MACHINE SCORING (FAQ)

What is the difference between machine scoring and generative AI?

Machine scoring evaluates language performance against predefined proficiency criteria. Generative AI creates content. Unlike generative AI, which predicts and produces language, machine scoring operates within tightly defined parameters. In assessments such as the Spanish AAPPL, scoring models are trained on validated responses aligned to recognized standards, including the ACTFL Proficiency Guidelines and ACTFL Performance Descriptors. The goal is not to improvise; it is to replicate trained human judgment consistently and at scale.

How does machine scoring learn to score speaking and writing?

Machine scoring models learn from large collections of authentic responses that have already been evaluated by ACTFL-certified human raters. They cannot simply be turned on and expected to rate writing or speaking accurately. The model learns the language features and performance characteristics associated with different proficiency levels and uses those patterns to evaluate new responses.

In that sense, machine scoring is not inventing scores; it is learning to mirror expert human ratings based on extensive evidence.

How do you know that the scores are accurate?

During development, machine-generated scores are compared with a validation dataset that is composed of test scores assigned by ACTFL-certified human raters and unseen by the machine scoring system. Statistical analyses are used to measure agreement and identify where the model performs well or needs improvement. The Spanish AAPPL machine scoring was validated against ACTFL-certified ratings over multiple years of administration. Ongoing monitoring and validation are necessary to confirm that the system continues to perform as intended over time.

How is fairness addressed?

Fairness begins with representative training data and continues through ongoing review, validation, and monitoring. ACTFL’s and LTI's approach is informed by established language assessment ethics principles and supported by independent expert oversight.

Can machine scoring work equally well in every language?

Only when there is enough high-quality, language-specific evidence behind it. Then, the machine needs to be trained on this language-specific evidence. Machine scoring depends on sufficient representative training data, validation against human raters, and ongoing monitoring. Without that evidence, scores may be less consistent, less accurate, or less fair.

Because testing volume varies by language, not every language has the sufficient data needed to support responsible machine scoring. Spanish AAPPL does, which is why machine scoring is currently available only for Spanish AAPPL.

The key principle is simple: efficiency should never come before evidence.

What happens if the machine score and human score disagree?

Disagreement between raters can happen in any scoring system, including systems that solely use human raters. If a disagreement occurs, an additional human rating is used to resolve the score. Human judgment remains the benchmark for evaluating machine scoring performance.

What questions should schools ask about AI scoring?

Schools should ask:

  • What data was used to train the system?
  • Was the data representative?
  • What proficiency standards were used?
  • Has the system been validated against human raters?
  • How is the machine’s performance monitored?
  • How often is the scoring model reviewed and recalibrated?

Trust should be based on evidence, not marketing claims.

See a complete list of FAQs

EVIDENCE MATTERS

Not all AI scoring solutions deserve the same level of confidence.

Choose a machine scoring approach built on recognized standards, human expertise, rigorous validation, ethical oversight, and years of research.

Voss, E., Sallee, K., Son, Y. A., Malone, M. E., Marshall, C., & Chomon-Zamora, C. (2026). An Automated Scoring System for the AAPPL Spanish Presentational Writing Tasks. Foreign Language Annals,59(1), 63-81.

Voss, E. (2025). Comparison of Traditional Machine Learning and Neural Network Approaches for Automated Scoring of Second Language English Essays. Language Testing, 42(4), 369-396.

Voss, E. (2024). Artificial Intelligence in Language Assessment. In A. J. Kunnan (Ed.), The Concise Companion to Language Assessment(pp. 112-125). Wiley-Blackwell.

Voss, E., Cushing, S. T., Ockey, G. J., & Yan, X. (2023). The Use of Assistive Technologies Including Generative AI by Test Takers in Language Assessment: A Debate of Theory and Practice. Language Assessment Quarterly, 20(4-5), 520-532.

Chapelle, C. A., & Voss, E. (2008). Utilizing Technology in Language Assessment. In N.H. Hornberger (Ed.), Encyclopedia of Language and Education(2nd ed., Vol. 7, Language Testing and Assessment, pp. 123-134). Springer.



Automated Scoring System for a Spanish Writing Test: Feature-Engineered Machine Learning vs. Large Language Model Approach.

2024 AAAL Paper presentation.

Erik Voss, Young-A Son. Houston, TX. March 16, 2024.

An Argument for an Automated Scoring System for Spanish Written Responses: Supporting an Evaluation Inference.

2023 ECOLT Conference Presentation.

Erik Voss, Young-A Son. Georgetown University, Washington, D.C., October 27, 2023.

An Operational Automated Scoring System for Written Constructed Responses in Spanish.

2023 LTRC Conference Presentation.

Erik Voss, Dario Dathaguy, David Luongo. New York City, June 9, 2023.

A Validity Argument to Support Automated Scoring Models for a Test of Written Spanish.

2022 LTRC Conference Presentation.

Erik Voss, Hung Phan & Xijia Wang. Language Testing Research Colloquium (LTRC), (Tokyo) online, March 7-8, 2022.

Automated Scoring of Written Constructed-Responses in Spanish.

2021 ECOLT Conference (Online) Presentation.

Erik Voss, Hung Phan & Xijia Wang. October 22, 2021.

Adapting an Automatic Speech Recognition (ASR) Model for Spanish Language Assessment.

2026 AAAL Conference Presentation.

Kim Sallee, Young-A Son & Meg Malone. Chicago, Illinois. March 22,2026.

Automated Topic Detection in Spoken Language Assessment: Comparing Human-Labeled Training and Prompt-Based Category Prediction.

LTRC 2026 Poster Presentation.

Son, Y-A., Voss, E., Sallee, K., and Malone, M. Montreal, Canada. June 2-6, 2026.

Affordances and Challenges of Developing an Automated Spoken Spanish Assessment.

2025 AAAL Conference Presentation.

Sallee, K., Voss, E., Son, Y-A., Malone, M. Denver, Co. 2025, March 22-25.

Artificial Intelligence, Automated Essay Scoring, Large Language Models (EALTA 2024).

2024 EALTA Conference Presentation.

Voss, E., Malone, M., Sallee, K. Belfast, Northern Ireland. June 4-9, 2024.

Applying the ILTA Code of Ethics to Responsible Language Testing.

2026 AAAL Presentation.

Malone, M, Deygers, B. Chicago, IL 2026.

Accountability and Ethics in Language Assessment: Examining Approaches to Machine Scoring.

2025 ECOLT Conference Presentation.

Malone, M., Sallee, K., Gravina, S., Son, Y-A. Washington, DC. September 2026.

Maintaining Ethical and Secure Data Practices in Machine Scoring of the AAPPL Test.

2025 AIRiAL Conference Presentation.

Gravina, S., Sallee, K., Sun, S. September 26-27, 2025.

Automated Scoring System for a Spanish Writing Test: Feature-Engineered Machine Learning vs. Large Language Model Approach.

2024 AAAL Paper Presentation.

Erik Voss, Young-A Son. Houston, TX. March 16, 2024.

Explore the complete AI Research Library & Resource Hub

Learn More About ACTFL & LTI Machine Scoring
Copyright © 2026 Language Testing International. All rights reserved.