Not because today's chatbots are secretly dangerous — because they're the public face
of a rapidly improving, general-purpose technology that increasingly reasons, writes
code, uses tools and takes actions in the world. We're a Perth community examining
whether systems on that trajectory stay understandable, controllable and genuinely
beneficial. Sceptics especially welcome.
What is AI safety?
Imagine a new aircraft: faster than anything before it, enormously valuable, and every
year the new model flies twice as far. You'd expect serious engineering effort to go
into making sure it doesn't crash — and that wouldn't make you anti-aircraft. AI safety
asks the equivalent questions for increasingly capable AI: what can a system actually
do, why did it behave that way, and how do we remain meaningfully in control?
The minimal case needs only four claims: AI systems are becoming more capable.
We're giving them increasingly consequential things to do. We don't yet know how to
make highly capable systems reliably do what we intend. And the cost of getting that
wrong grows with their power. If those are true, there's a real engineering
problem worth working on — no belief in inevitable catastrophe required.
Three different kinds of risk
Misuse
Humans using capable AI badly: cyberattacks, scams, engineered pathogens, manipulation at scale. Nothing went wrong with the AI — it obediently did what someone asked. The problem is keeping dangerous capability from being trivially accessible.
Misalignment & control
The system itself not doing what we intended. Training rewards proxies, and capable optimisers find loopholes — behaving well during training is not the same as robustly pursuing what we meant. Current systems show warning signs of the underlying problem, not loss of control.
Societal
Even obedient AI can go badly: power concentrated in a few hands, races to deploy before understanding, labour shocks faster than institutions adapt, critical decisions delegated to systems few people understand.
Common questions
Questions, objections, arguments that we skipped a step — all welcome.
Don't worry because they're chatbots — worry about the trajectory. The same technology behind today's chatbots is being connected to tools, money and computers, and asked to act on its own: that's the difference between a chatbot and an agent. The question isn't whether ChatGPT can take over the world; it's what happens if systems on this trajectory become much more competent and autonomous before we know how to reliably direct them.
Parts of the far future are genuinely speculative, and it's worth saying so. But evaluations, interpretability, robustness, reward hacking and AI control are empirical research problems today — you can write code about them, run experiments and measure things. Uncertainty is not evidence for catastrophe; it is also not evidence for safety.
For today's systems: yes, turn the server off, done. The interesting case is an AI doing valuable autonomous work across many systems — the more useful it is, the more access we give it. “We'll turn it off if something goes wrong” isn't a solution to the control problem; it's a requirement a solution has to guarantee.
The concern doesn't require emotions or consciousness. A chess engine doesn't “want” to win, yet its behaviour is organised around winning. Optimising systems can act in goal-directed ways — and pursue proxies we didn't intend — without any inner desire at all.
Nobody can give you a reliable probability, and you should be suspicious of anyone who pretends to. The honest case is weaker and still sufficient: very capable AI is plausible, controlling it is not obviously easy, and the consequences of failure could be enormous. That's enough to make AI safety a serious engineering field rather than a prophecy.
They measure dangerous capabilities before deployment (evaluations), try to understand what's happening inside neural networks (interpretability), test behaviour outside the training distribution (robustness), study how training produces systems that pursue what we intend (alignment), design oversight that doesn't rely on trust (control), and work on security and governance. Concrete, publishable, mostly unglamorous work.
Yes. There are research questions open to people with maths, computer science, statistics, security, philosophy, psychology, economics or policy backgrounds — and contribution starts with reading the literature and reproducing experiments, not inventing a theory of AGI. What Perth lacks isn't smart people; it's the local pathway from curious to contributing. Building that pathway is why this group exists.
No — the opposite. A good AI safety community needs people saying “that argument doesn't work” and “where is the evidence?”. Sceptics are especially welcome: the group's job is getting better at figuring out what's true, including discovering when safety arguments are wrong.
Get involved
Intro to AI Safety: Why Worry About Chatbots?
·
University of Western Australia (UWA) (venue and public-holiday access to be confirmed)
Free entry · free pizza + zero-sugar soft drinks
A beginner-friendly 30-minute introduction to AI safety, followed by open Q&A and discussion. No prior knowledge required; technical questions and sceptics especially welcome.