Thursday — May 28, 2026

Anthropic just got backup from the enemy camp.

Build an AI Ethics Policy

Act as a legal and ethics advisor for an AI startup. Draft a 3-paragraph internal policy outlining red lines for military or defense applications, including employee opt-out rights and public transparency commitments. Use concise, enforceable language.

01

Google and OpenAI employees back Anthropic in Pentagon fight

Google and OpenAI employees back Anthropic in Pentagon fight

On Thursday, the AI industry witnessed an unprecedented show of force — not from corporate giants, but from their own employees. Over 300 current and former workers from Google, OpenAI, and other leading AI labs filed an amicus brief in support of Anthropic, which is locked in a legal battle with the U.S. Department of Defense over refusing to deploy its AI for specific military purposes. The brief argues that AI companies should have the right to refuse contracts that violate their ethical guidelines, and that employees should not be forced to build systems they believe could cause harm.

This is not an abstract debate. The Pentagon has been aggressively pursuing AI partnerships for everything from logistics to autonomous systems, and Anthropic's refusal to play ball made it a target. The lawsuit could set a precedent for whether AI companies can pick and choose their government clients based on ethical red lines. If the court rules against Anthropic, it could force every AI lab to accept military contracts or face legal consequences — a chilling prospect for builders who care about responsible AI.

What makes this story wild is the cross-company solidarity. Google and OpenAI are direct competitors with Anthropic in the AI arms race, yet their workers are publicly backing a rival. This suggests that the grassroots ethical movement in AI is stronger than corporate loyalty. For builders, this is a signal that your values matter more than your employer's stock price — and that the talent market will increasingly reward companies with clear ethical principles.

The immediate practical implication: if you're building AI products today, you need to define your own red lines before someone else defines them for you. Whether it's military use, surveillance, or content moderation, the legal and reputational risks of being reactive are enormous. Anthropic is taking the hit now, but every builder will feel the ripple effects of this case.

For those working on AI safety or alignment, this is also a reminder that technical solutions alone won't save you. The ethical stance you take publicly will determine whether top talent wants to work with you — and whether your users trust you. The amicus brief includes signatures from researchers who worked on GPT-4, Gemini, and Claude, proving that even the architects of these systems have limits.

Finally, this story is a wake-up call for founders and engineering leads. If you haven't had a team-wide conversation about what you will and won't build, do it this week. Your employees are watching, and they might just side with your competitor's team if you don't lead with integrity.

Flo's take: Holy crap, when your biggest rivals' workers show up to court for you, you know you've struck a nerve. This is the AI industry's 'we won't build that' moment — and it's happening in real time.


02

AI talent wars have come for the interns

AI talent wars have come for the interns

The AI talent war has officially hit the college campus. According to a Business Insider report published today, companies like OpenAI, Anthropic, and Google are offering intern salaries that rival mid-level engineering roles at non-AI companies. We're talking base salaries of $10,000 to $15,000 per month, plus signing bonuses, relocation packages, and guaranteed return offers. For a summer intern, that's a six-figure annualized comp — unheard of in any previous tech cycle.

Why the frenzy? Because AI is moving so fast that even junior talent with a few months of experience building with LLMs or fine-tuning models is considered a hot commodity. Companies are betting that today's intern is tomorrow's lead researcher, and they're willing to pay a premium to lock that relationship in early. The strategy is simple: get them in the door, show them the culture, and make it painful to leave by the time they graduate.

For builders, this changes the calculus of how you staff your AI projects. If you're a startup, you can't compete on salary alone — but you can compete on impact and ownership. Interns at big labs often get relegated to narrow tasks, while a startup can give them end-to-end ownership of a model deployment or a critical feature. That's a powerful recruiting tool.

The downside? This arms race is inflating compensation expectations across the board, making it harder for smaller teams to hire any AI talent at all. If you're bootstrapping, you might need to rethink your hiring strategy entirely — focus on remote talent, non-traditional backgrounds, or even hiring from adjacent fields like computational linguistics or data engineering.

There's also a cultural risk. Interns who are treated like royalty from day one may develop unrealistic expectations about career progression. But for now, the market is so hot that most companies are willing to take that gamble. The key takeaway for builders: start building relationships with universities and student AI clubs now. The best interns are being snapped up before they even submit their first application.

Finally, this trend underscores the importance of brand. Interns are choosing companies based on mission and ethics as much as comp — which is why Anthropic's principled stance on military AI might actually help it win the talent war, even if it can't match Google's salary bands. Builders should lean into their unique culture and values to attract the next generation of AI talent.

Flo's take: Interns making more than most senior devs? Wild times. If you're not already poaching undergrads from top CS programs, you're losing the talent race before it even starts.


03

Google, OpenAI, Anthropic race their AIs on Pokémon

Google, OpenAI, Anthropic race their AIs on Pokémon

In what might be the most entertaining AI benchmark of 2026, Google, OpenAI, and Anthropic are pitting their flagship models against each other in a Pokémon Red/Blue speedrun competition. Each model controls a separate Twitch stream, making decisions in real-time as they navigate Kanto, catch Pokémon, battle gym leaders, and attempt to beat the Elite Four. The streams have become a spectator sport, with thousands of viewers watching their favorite AI stumble, learn, and occasionally pull off brilliant strategies.

Why Pokémon? Because it's a surprisingly rigorous test of an AI's reasoning abilities. The game requires long-term planning (which gym to tackle next, which Pokémon to train), short-term tactics (which move to use in battle), and the ability to recover from failure (when your starter faints or you get lost in a cave). It's a closed-world environment with clear rules, but the combinatorial complexity is enormous — there are hundreds of Pokémon, moves, and items to consider.

Early results show that each model has distinct strengths and weaknesses. OpenAI's model is aggressive, often overleveling its starter and brute-forcing battles, but it struggles with puzzles like the Safari Zone. Google's Gemini is methodical, exploring every route and item, but it can get stuck in analysis paralysis. Anthropic's Claude is the most cautious, often avoiding risky battles and prioritizing healing, but it sometimes misses opportunities for big gains.

For builders, this is more than a fun distraction. It's a live demonstration of how different architectures and training approaches handle real-time decision-making under uncertainty. If you're building an AI agent that needs to operate in a dynamic environment — like a customer support bot or a trading algorithm — watching these streams can give you intuition about which model behavior you prefer.

There's also a practical lesson: benchmarking matters, but creative benchmarks like this reveal failure modes that standard tests miss. A model that scores 99% on a multiple-choice reasoning test might still get stuck in Viridian Forest for hours. If you're evaluating AI for your own use case, don't rely solely on leaderboards — design a test that mimics your actual deployment conditions.

Finally, this competition highlights the growing trend of gamification in AI research. By making benchmarks public and entertaining, companies can engage the community, attract talent, and get free feedback on their models' weaknesses. Builders should consider running their own gamified tests — even a simple game like Pokémon can surface issues that would never appear in a controlled lab setting.

Flo's take: This is the most fun AI research I've seen all year. Watching Claude get lost in Mt. Moon is oddly relatable. But seriously, if you're building agents, this is a masterclass in why real-world testing beats synthetic benchmarks every time.

Deep Dive

How to Build an AI Agent That Plays Pokémon (and Why It Will Make You a Better Builder)

You don't need to train a foundation model to build an agent that can play Pokémon. In fact, you can do it with an API call to any major LLM and about 200 lines of Python. The trick is designing the right interface between the model and the game. Here's a step-by-step approach that I've used to get Claude, GPT-4, and Gemini all playing Pokémon Red on a local emulator.

First, you need a Game Boy emulator with a scripting interface. PyBoy is the gold standard here — it's a Python-based Game Boy emulator that gives you pixel-level access to the screen and memory. You can read the current frame, detect sprites, and even read the game's internal variables like the player's position, Pokémon party, and item inventory. This is your AI's eyes and ears.

Next, you need to translate that raw game state into a text prompt. Don't feed the model raw pixels — it's too slow and expensive. Instead, extract key information: your current location, your Pokémon and their HP, the wild Pokémon you're facing, and your available items. Format this as a structured text block. For example: 'Location: Route 1. Party: Charmander (HP 20/20), no other Pokémon. Wild encounter: Pidgey (HP 12/12). Available actions: FIGHT (Scratch, Growl), BAG (Potion x3), RUN.' The model can then reason about what to do.

Now, the hard part: prompting. You need to give the model a persona that encourages strategic thinking. I use a system prompt like: 'You are an expert Pokémon trainer. Your goal is to beat the Elite Four as fast as possible. You must optimize for speed, not completion. Avoid unnecessary battles. Prioritize leveling your starter. Use items sparingly. Output only a single action: FIGHT [move], BAG [item], SWITCH [Pokémon], or RUN.' This constrains the output to a parseable format while giving the model enough context to make smart decisions.

But here's where most builders fail: they don't handle errors. The model will occasionally output an invalid action — like trying to use a move that doesn't exist or switching to a fainted Pokémon. You need a robust error handler that either retries with a corrected prompt or falls back to a safe action like 'RUN'. I log every invalid action and feed it back into the next prompt as a lesson: 'Your previous action was invalid because Charmander does not know Water Gun. Choose a valid action.' This gives the model a chance to learn from mistakes.

Finally, you need to think about speed. LLM API calls take 1-3 seconds, but the game keeps running. You can either pause the emulator while waiting for a response (which makes it less realistic) or let the game run in real-time and have the AI react every few frames. I prefer the latter — it forces the model to make decisions under time pressure, which is closer to real-world agent deployment. You can also cache common responses for repetitive situations like wild battles.

The result? A surprisingly competent AI player that can beat the first few gyms without much trouble. The real value, though, is what you learn about prompt engineering, error handling, and state representation. These skills transfer directly to building any AI agent — whether it's a chatbot, a code assistant, or a trading bot. So fire up PyBoy, grab an API key, and start training your own Pokémon master. Your users will thank you.

Build red lines, not just models — your team is watching.

Get this in your inbox

Real AI news every morning. No fluff. Free.