UC-UCO-CICESE_UT3-Plenitas Team - Exploring in the PROFE 2025: Language Proficiency Evaluation
Yoan Martínez‐López, Mayte Guerra Saborit, Yanaima Jauriga, Mercedes Leguen De Varona, Julio Madera, Ansel Y. Rodríguez‐González, Carlos de Castro Lozano, José Miguel Ramírez Uceda, José Carlos Arévalo Fernández · 2025
PROFE2025 is a challenge that tests automated systems on authentic Spanish reading-comprehension exams used by the Instituto Cervantes for human learners, without providing task-specific training data. This setup emphasizes transfer learning and large generative models. Participants can choose from three subtasks: Multiple Choice, Matching, and Fill-in-the-Gap, each evaluated using accuracy metrics that align with human assessment. Teams completing all subtasks receive an exam-level score based on the proportion of exams passed (60% accuracy), enabling direct comparison to human performance. The competition utilizes the IC-UNED-RC-ES corpus and Hugging Face's transformer model ecosystem. In internal evaluations, three Qwen-3 variants were tested: Qwen 3.0, a "think" prompting version, and a DeepSeek-distilled Qwen 2.5 7B. Both base and distilled models achieved 32.22% overall accuracy, while the "think" variant scored only 10.19%, indicating room for improvement in reasoning-focused prompts. Among 40 submissions, only two teams completed all three subtasks. UC-CICESE_UT3-Plenitas placed second with 32.22%, 4.89%, and 10.19%, respectively, underscoring both potential and current limitations of LLMs in this domain. PROFE2025 highlights that large language models can engage meaningfully with complex, human-like exam conditions, but significant improvements are possible through specialized adaptation, fine-tuning, and advanced prompting strategies, positioning it as a key benchmark for progress in automated language understanding..