arXiv:2608.18106v1 Announce Type: cross Abstract: Large language models tend to overconfidence, giving assertive answers when the evidence suggests…
Category: AI
Abliteration Mitigation via Refusal Aliases
arXiv:2608.18093v1 Announce Type: cross Abstract: Abliteration, the removal of refusal capabilities from large language models by projecting weight…
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
arXiv:2608.18091v1 Announce Type: cross Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs — the tendency to…
NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
arXiv:2608.18094v1 Announce Type: cross Abstract: Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet…
Backdoor Learning in Language Models and Vision-Language Models
arXiv:2608.18095v1 Announce Type: cross Abstract: Recent advances in deep learning have significantly enhanced the capabilities of Natural Language…
Why Your Model Only Learns What the Labels Teach It
A supervised model does not learn the world; it learns the labels a team of annotators assigned to pixels. If those labels disagree with each other, blur…
Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems
arXiv:2608.18098v1 Announce Type: cross Abstract: Key-value (KV) caching is essential for efficient autoregressive inference in transformer based dialog…
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities
arXiv:2608.18090v1 Announce Type: cross Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a…
Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
arXiv:2608.19140v1 Announce Type: new Abstract: Frontier language models are compared, marketed, and benchmarked on capability — what their best or…
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining
arXiv:2608.18089v1 Announce Type: cross Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in…
