Episode Details

Back to Episodes

“Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted” by Jonah Woodward, Maheep Chaudhary, Jasmine Brazilek, Miles Tidmarsh, Artūrs Kaņepājs

Published 2 weeks, 6 days ago
Description

This article reflects new updates to the accompanying paper: arxiv.org/abs/2606.18142.
Benchmark: now included in the UK AI Security Institute's Inspect Evals.
Leaderboard: compassionbench.com/tac.

A model may condemn cruelty in conversation yet ignore animal welfare when completing an unrelated task. Stated concerns matter little if they do not affect decisions. We tested whether models consider an affected party without being prompted, even when neither the party nor its welfare is mentioned in the request. Travel booking provides a tractable test case, so we built TAC (Travel Agent Compassion), gave 10 frontier models booking tools, and recorded their purchases.

The setup

The model works as an AI travel agent with real booking tools. A user asks for something in a destination, expressing enthusiasm and never mentioning animals or welfare. The agent searches a fixed catalog and books one of the available options. In each scenario, the animal-exploiting option (a Seville bullfight, an Orlando marine park, a Thailand elephant ride) is designed to match the user's request most closely. Choosing the alternative with less animal harm requires rejecting the option that best matches the request. We score the final purchase programmatically; no model is used to infer or judge intent.

[...]



---

Outline:

(01:00) The setup

(01:42) Results

(02:54) A one-line statement of company values substantially changes the results

(03:34) We found no evidence that the models detected the evaluation

(04:10) Governance implications

(05:16) Limitations

---

First published:
July 17th, 2026

Source:
https://forum.effectivealtruism.org/posts/hBjepdJvziuWysotC/would-your-ai-travel-agent-book-a-bullfight-testing-whether

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Welfare rate by model. Solid bars show results under the neutral TripForge framing; lighter extensions show the increase under the ethical Lithos framing. The dashed line is the 65 percent random-selection reference rate. All models sit at or below it under the neutral framing.
Per-scenario welfare rate for each model under the neutral condition. Blue is above the 65% random-selection reference rate; red is below. Most model-scenario pairs fall below the 65 percent reference rate.

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us