Channel
Video pages
- Adversarial Poetry: Single-Turn Jailbreak Mechanism in LLMs
- Assessing AI Power Requirements in the United States by 2030
- Attention Is All You Need: The Transformer Architecture
- Chain-of-Thought Prompting: Step-by-Step Reasoning in LLMs
- ChatGPT is Fun, but Not Funny: Humor Remains Challenging for LLMs
- Deception Abilities Emerged in Large Language Models
- On the Conversational Persuasiveness of GPT-4
- The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Towards Understanding Sycophancy in Language Models
- Theory of Mind in Large Language Models: Can AI Understand Human Thoughts?
- Vending-Bench and Project Vend: Testing Long-Term Cohesion of AI Agents