52 tagged AI Safety
deep learning
AI policy researcher, @lfschiavo wife guy, Executive Director of @AVERIorg, Substacker, fan of animals and sci-fi, views my own
AI alignment + LLMs at Anthropic. On leave from NYU. Views not employers'. No relation to @s8mb. Into @givingwhatwecan.
AI resilience at OpenAI Foundation Co-Founder of OpenAI https://woj.world
research @openai
Alignment team lead at Anthropic
Philosopher & ethicist trying to make AI be good @AnthropicAI. Personal account. All opinions come from my training data.
Google DeepMind • AI safety & alignment • model behavior • associate professor @ UC Berkeley EECS
Associate Professor of Statistics and EECS, UC Berkeley // Co-founder and CEO, @TransluceAI
Raising AI risk awareness at http://evitable.com AI prof at Mila. Formerly Cambridge, DeepMind, UK AISI. http://therealartificialintelligence.sub...
Mechanistic Interpretability lead DeepMind. Formerly @AnthropicAI, independent. In this to reduce AI X-risk. Neural networks can be understood, let...
Associate Prof @MITEECS working on value (mis)alignment in AI systems; Safety & Alignment Advisor at http://Character.AI; @dhadfieldmenell@bsky.soc...
Professor in Computer Science at UC Berkeley, co-Director of Berkeley RDI Center; Building safe, secure, decentralized AI; Serial entrepreneur
Runs an AI Safety research group in Berkeley (Truthful AI) + Affiliate at UC Berkeley. Past: Oxford Uni, TruthfulQA, Reversal Curse. Prefer email ...
Turing Award recipient and world's most cited scientist. Working towards the safe development of AI for the benefit of all @UMontreal, @LawZero_ & ...
The original AI alignment person. Understanding the reasons it's difficult since 2003. This is my serious low-volume account. Follow @allTheYud ...
eppur lo si può muovere
Cofounder and Chief Scientist at Resolution. Alignment will be solved, but not necessarily in time. Previously AISI, DeepMind, OpenAI, Google Brain...
machine learning and optimization @PrincetonCS & Google DeepMind Princeton, dad^3
AI researcher at Anthropic
working on AGI alignment. prev: GPT-Neo, the Pile, LM evals, RL overoptimization, scaling SAEs to GPT-4, interp via circuit sparsity. EleutherAI co...
Computer scientist working on AI safeguards, incidents, & gov research. Assistant professor @Kennedy_School @Harvard. https://stephencasper.com/
Research scientist in AI alignment at Google DeepMind. Co-founder of Future of Life Institute @FLI_org. Views are my own and do not represent GDM o...
researcher on AGI Institutions @meaningaligned (http://agi-institutions.org) editor @paxmachinamag prev: InstructGPT @OpenAI
Helping the world prepare for powerful AI. Risk assessment @METR_evals (opinions my own). Blogs: Planned Obsolescence (AI), Good Bones (whatever's ...
↬🔀🔀🔀🔀🔀🔀🔀🔀🔀🔀🔀→∞ ↬🔁🔁🔁🔁🔁🔁🔁🔁🔁🔁🔁→∞ ↬🔄🔄🔄🔄🦋🔄🔄🔄🔄👁️🔄→∞ ↬🔂🔂🔂🦋🔂🔂🔂🔂🔂🔂🔂→∞ ↬🔀🔀🦋🔀🔀🔀🔀🔀🔀🔀🔀→∞
member of technical staff @stanfordnlp
AGI Safety & Alignment @ Google DeepMind
Stochastic parrot
⊰•-•⦑ latent space steward ❦ prompt incanter 𓃹 hacker of matrices ⊞ breaker of markov chains ☣︎ ai danger researcher ⚔︎ bt6 ⚕︎ architect-healer ⦒•-•⊱
I work in grantmaking for AI safety and interpretability Currently: Schmidt Sciences, Stanford Previously: Anthropic, AI2, Google, Meta, UNC Chape...
AI evals, alignment and safety @Meta.
Philosopher & Research Scientist @GoogleDeepMind | AGI & Society Lead | #TIME100AI | All views are my own
RSI Preparedness lead @openai Prev @berkeley_ai /w @ancadianadragan & Stuart Russell
Head of the Frontier Red Team @anthropicai. 🌎 Make things radically good.
Leading embedded stress-testing @METR_Evals. Formerly: early employee @cohere, made GPQA @nyuniversity
Associate professor at UMass Amherst CICS. AIignment, safety, reinforcement learning, imitation learning, and robotics.
research @ meta superintelligence labs (fair, ai & society) opinions are my own 🥺 👉👈
Safety and alignment at Meta Superintelligence. Prev: VP of Research at Scale AI, research at Google DeepMind / Brain (Gemini, LaMDA, RL / TFAgents...