28What are your views on AI safety and risk? What would you refuse to build for a customer?▼hard★ EssentialAnthropicOpenAI1 replies◆ premiumAnthropic's signature values filter, and it screens out two groups: candidates with no real views, and candidates performing views they clearly downloaded last night. Here's what an authentic, FDE-grade answer contains.Open full answer →
42Prove your financial-advice LLM has no internal 'deceptive' policy: use SAEs to find, validate, and suppress deception features.▼expertAnthropicGoogle DeepMindOpenAI1 replies◆ premiumSAEs can surface candidate 'deception' features in the residual stream, but a correlated feature is not a cause. The real work, and the honest answer, is causal validation and admitting what interpretability cannot yet prove.Open full answer →