ZICQ
中 Log in / Sign up
Newsroom Research & Papers #Language Models #Brain Alignment #Cross-Lingual Transfer #Scoring Reliability #ArXiv

Study Reveals: Unreliability of Scores in Language Model Brain Alignment and Cross-Lingual Transfer

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:ArXiv has released a study on the alignment of language models with the human brain and cross-lingual transfer, highlighting critical flaws in traditional similarity score-based evaluation methods. The research shows that when shared structures are absent or measurement tools fail, the scores become unreliable. For instance, grammatical probes perform worse with increasing language distance, rendering the scoring tool ineffective in 64 out of 272 language pairs. Furthermore, only a small portion


Background and Motivation

In recent years, similarity scores between language models and brain signals or cross-lingual structures have been widely used as key indicators to evaluate model performance and alignment. However, the validity of these scoring methods has lacked in-depth exploration. This study aims to reveal the limitations of these scoring methods under specific conditions and propose improvements.

Key Findings

  1. Failure of Scores in Cross-Lingual Transfer

    • The study found that probe models trained to distinguish between grammatical and ungrammatical sentences perform poorly in cross-lingual transfer, especially for language pairs with greater linguistic distance.
    • In 17 languages, the probe's performance was at chance level for 4 languages, rendering the scoring tool ineffective for 64 out of 272 language pairs.
    • Excluding these language pairs halved the strength of the relationship, indicating that the scoring tool fails for language pairs with greater linguistic distance.
  2. Issues with Brain Alignment Scores

    • Training a language model to align with human brain signals can increase the similarity score from 0.10 to 0.34, but this is still far from the ceiling of 0.54.
    • Even when the target's correspondence to the brain was destroyed, the model still scored 0.31, suggesting that only 0.028 to 0.068 of the increase is attributable to the brain itself.
    • Without any model, the destroyed target already scores 0.204 from the real one, indicating that the score is heavily influenced by the target's rank and the number of sentences.
  3. Evaluation of Language-Guided Directions

    • In 17 languages, 16 showed that guiding along their own direction was more effective than a random direction (average increase of 6.2), but this improvement was not correlated with language distance.

Conclusions and Recommendations

This study reveals the limitations of similarity score-based evaluation methods in assessing the alignment of language models with the brain and cross-lingual transfer. The researchers recommend more cautious interpretation of alignment scores and the exploration of more reliable evaluation methods, such as combining multimodal data, introducing causal reasoning mechanisms, or developing new evaluation metrics.

Recommendations for Developers

  • Use Scoring Indicators Cautiously: When evaluating model alignment with the brain or cross-lingual structures, combine multiple evaluation methods and avoid over-reliance on a single scoring indicator.
  • Explore New Evaluation Methods: Developers can experiment with causal reasoning, multimodal data fusion, and other new technologies to improve evaluation reliability.
  • Focus on Model Robustness: In cross-lingual transfer tasks, pay attention to model robustness and optimize for different linguistic distances.

Source: ArXiv NLP/LLM (cs.CL) (2026-10-06)

— END —

Tags: #Language Models #Brain Alignment #Cross-Lingual Transfer #Scoring Reliability #ArXiv

Community Comments

Loading live comments and annotations…