Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
Abstract
Abstract Large Language Models (LLMs) exhibit position bias systematically underweightinginformation based on its location in the context but how this bias varies acrosslanguages and models remains unclear. We conduct a multilingual study across fivetypologically diverse languages (English, Russian, German, Hindi, Vietnamese) andfive model architectures, analyzing how position bias interacts with promptingstrategies and affects output entropy. Our key findings are: (1) Position bias isprimarily model-driven but shows language-specific nuances. Notably,Qwen2.5-7B-Instruct, DeepSeek 7B Chat and Mistral 7B consistently favor latepositions challenging the common assumption of universal early-token preference.(2) Explicit relevance scoring that “the most relevant context to the query ismarked as 1 ” consistently reduces accuracy regardless of scoring correctness, withscore omission yielding the strongest performance across nearly all settings. (3)Accuracy consistently drops most when relevant information appears in the middleof the context, yet this is not reflected in a corresponding increase in outputentropy, suggesting the model remains confident even when it fails to usemid-context cues.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.