arXiv:2608.17202v1 Announce Type: new Abstract: Safety alignment in open-weight language models is trivially removable: abliteration projects a…
Category: AI
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive…
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
arXiv:2608.17150v1 Announce Type: new Abstract: To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must…
A Constitution for the New Enterprise: 15 Rules for AI Governance
These essays were supposed to be about architecture. Read them back and notice what they actually did: they kept issuing rules. Approvals are not data;…
Toward Personal Intelligence Through Cooperative Observation
arXiv:2608.17128v1 Announce Type: new Abstract: A personal AI system needs a model of the user’s goals, constraints, and ongoing commitments to plan and…
Copilot Autofix Opened a Shell Injection in Snowflake’s CI/CD Pipeline
A security fix written by GitHub’s Copilot Autofix and merged into a Snowflake repository on June 18, 2026 stripped out a sanitized input pattern and left…
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection
arXiv:2608.17170v1 Announce Type: new Abstract: Algorithm selection for constraint satisfaction problems requires extracting features that capture problem…
A decodability criterion predicts when hidden-state selection beats majority voting in large language models
arXiv:2608.17124v1 Announce Type: new Abstract: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time…
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
arXiv:2608.17067v1 Announce Type: new Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the…
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools
arXiv:2608.17007v1 Announce Type: new Abstract: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate…
