Defining and Evaluating Physical Safety for Large Language Models
Publication signal: Defining and Evaluating Physical Safety for Large Language Models
Science community intelligence
AI league updates across foundation models, evaluations, embodied systems, AI for science, compute, and open tooling.
Publication signal: Defining and Evaluating Physical Safety for Large Language Models
Publication signal: Ground-truthing AI energy consumption: validating CodeCarbon against external measurements
Publication signal: PsyEval: a comprehensive large language model evaluation benchmark for mental health
Publication signal: From reactive filtering to proactive moral architecture: rethinking ethical alignment in large language models
Publication signal: Can AI help reduce prejudice? Evaluating the effectiveness of AI-powered personalized persuasion on support for transgender rights
Publication signal: Explanations of the Fermi Paradox and the Drake Equation
Publication signal: AI, Digital Platforms, and the New Systemic Risk
Publication signal: The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
Publication signal: Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power
Publication signal: AI Alignment and Safety of Large Language Models: A Survey of RLHF, Constitutional AI, Red-Teaming, and Value Learning
Publication signal: AI Alignment and Safety of Large Language Models: A Survey of RLHF, Constitutional AI, Red-Teaming, and Value Learning
Publication signal: Off-Centre AI
Publication signal: Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Publication signal: A Prompt-Based AI Safety Evaluation of Gender Bias in Large Language Models: A Comparative Study of ChatGPT and Gemini
Publication signal: A Prompt-Based AI Safety Evaluation of Gender Bias in Large Language Models: A Comparative Study of ChatGPT and Gemini
Publication signal: Compliance without coherence: fluent failure and the ethics of alignment evaluation in multi-agent language models
Publication signal: BEADS: Bias Evaluation Across Domains
Publication signal: What Does Your AI Mean by "Flourishing"? The Case for Disclosing and Benchmarking the Values in AI Alignment (Preprint)
Publication signal: Motivation and post-design evaluations of AI usage behind AI-assisted design
Publication signal: PRINCIPLED FRAMEWORKS FOR AI ALIGNMENT: FROM POST-TRAINING TO INFERENCE