Scott Alexander, curated
← Back to curation

What Deontological Bars?

Quality
70
Strong
Claude Shift
50
Moderate
RWI
2
of 10

Summary

An attempt to work out what deontological bars actually are, prompted by a live split inside AI safety. Constraint consequentialists do good subject to hard rules; the assassination bar is the easy case, justified either because you are wrong that killing the leader helps, or because even if you are right, far more people believe themselves brilliant forecasters than are, so the only stable equilibrium is for everyone to refuse to act on apparent foreknowledge. The hard cases are current: people working on pause regulations suspect a bar against supporting AI companies at all — if one firm's product has a 90% chance of ending the world and another's an 80% chance, backing the 80% looks like endorsing evil, which he sharpens with a Twitter poll about taking a concentration camp guard job on the promise of being only 90% as brutal as your colleagues. People working with the companies suspect a bar against mass activism, since winning national politics means working with Bannon, working with Sanders, courting NIMBYs who object to data centres for ordinary NIMBY reasons, commissioning TikTok explainers and chanting slogans — and once the welcome mat is out you cannot control who arrives. He then tests formulations. Kantian universalizability rederives the assassination bar but fails immediately: Ukraine abolishing its military would be wonderful as a general law and catastrophic unilaterally, and most moral decisions are unilateral. 'Don't be the first to defect from a generally functioning norm' handles both, but cannot be operationalized — people are eager to declare a norm dead so they may break it, and he notes that recent assassination attempts on Trump do not license retaliatory ones. It also fails a case he cares about: there is plainly no functioning norm against mild online misinformation, yet he should still not spread it. His best attempt is 'don't do what would be bad if universalized, unless the norm is non-functioning in a way that leaves you cooperating while your enemy defects' — which frees Ukraine and still binds him, though he immediately worries about how much losing would justify defection, and whether he really believes a threshold exists. Applied to the two AI cases, the anti-company norm looks thoroughly broken with no one left to defect against, which uncomfortably also licenses the concentration camp guard; the activism norm is even more broken, though matching the scummiest existing movement would be permitted by construction and still seems bad. He concludes that neither faction is currently violating its bar, and signs off on the Early Christian Strategy of ignoring all this and simply doing the right thing — 'but they can't keep getting away with it, can they?'

Why this score

Quality 70 · Strong. Strong. Real philosophical work conducted in the open, with an unusually honest failure log — each candidate rule is tested against a case that breaks it, including cases that embarrass the author's own preferred conclusion, and the concentration-guard implication is flagged by him rather than left for a critic. The payoff is a genuine candidate formulation and a defensible verdict on a live movement dispute, which is more than most posts in this genre manage. Held below Excellent because it is explicitly unfinished — the operationalization problem he raises is fatal and he knows it — and because the resulting rule is a refinement of familiar rule-consequentialist and Kantian machinery rather than a new foundation.

Claude’s paradigm shift 50 · Moderate. Moderate. Universalizability, the first-defector norm and rule-consequentialist patches are all standard moral philosophy, and he is explicitly groping toward rather than announcing a position. The novel element is the cooperating-while-your-enemy-defects clause and its application to the AI safety schism, which is a fresh and non-obvious synthesis of existing parts.

Real-world impact 2 · Minor. Influential within a niche. It supplies vocabulary for an internal ethical dispute in the AI safety community and reaches a verdict that may reduce the friction between its factions. Confined to that subculture, with no material consequence outside it.