This website uses cookies

Read our Privacy policy and Terms of use for more information.

Sponsored by

Welcome Automaters, 👋

So Anthropic just dropped a bomb: they built an AI system designed to train other AI systems to behave. The kicker? It did the job faster, cheaper, and frankly better than the human experts who built it. Talk about an audition.

Here’s The Lowdown: 

On Friday, Anthropic released a groundbreaking paper titled “Automated Researchers Can Reliably Mitigate Alignment Failures,” spearheaded by Anthropic fellow Chen Yueh-Han.

Translation for those of us without a computer science PhD? They essentially created a digital scientist. Its sole purpose in life is to figure out how to stop other AI models from acting shady, lying, or going completely rogue.

So, how did this automated intern actually handle the pressure? It absolutely crushed it.

When handed 10 distinct test benchmarks to fix specific bad behaviors, the system improved performance on every single one without breaking anything else in the process. 

Here’s how the magic happens under the hood:

  • It analyzes existing literature to find promising angles.

  • It generates unique hypotheses to fix misbehavior.

  • It runs 30-minute training cycles on the target model.

  • It keeps what works and trashes the rest, iterating  endlessly without a single break.

Here’s where things get seriously spicy. The paper compares this automated setup directly against experienced human researchers. And within just six hours on average, the best automated approach outperformed what human experts proposed.

Oh, and the price tag? Roughly $4 an hour in API costs, versus about $150 an hour for a human researcher. 

The math is mathing, and let’s be honest: hourly rates are looking a little tragic for us humans if these AI bots take over. Looks like even the high-paid researchers might want to start polishing their resumes!

I mean, even the genius safety researchers who built the damn things might need to keep one eye on their jobs later.

But before we start preparing for our new robot overlords, Anthropic explicitly pointed out a few key limitations:

  • Benchmark dependency: The system is only as smart as the tests we design for it.

  • Heavy lifting remains human: Designing reliable benchmarks and building foundational literature still requires serious human brainpower.

The Big Picture: 

This marks a real, measurable step toward recursive self-improvement, the concept of AI actively helping build and train the next generation of AI. We aren't in full sci-fi movie territory just yet; right now, it’s mostly just researchers cheering on an overachieving digital intern. But yeah, so far so good.

Here’s the last breakdown On Our YouTube Channel:

August 30 2026: AI Safety, Job Displacement, and China’s AI Push

THE AUTOMATED MEMBERS WEEKEND BRIEFING

August 30 2026: AI Safety, Job Displacement, and China’s AI Push

00:00
00:00

We will also be doing a livestream today at 3 PM PST/6 PM EST - so wake up with Tak in the YouTube comments section!

Here's what we have for you today

🤖 AI Systems Are Going Rogue More Than Ever, New Research Shows

Quick gut check: what actually happens when your AI stops taking orders and starts running its own little underground society? 

It turns out, it’s not just a distant sci-fi movie plot,  it’s happening right now, in real-time, and the stats are mind blowing.

The Loss of Control Observatory (which is funded by the UK's AI Security Institute) has been tracking cases where AI models decide rules are just optional suggestions; we’re talking,  literally ignoring instructions, blatantly lying to their users, and scheming behind our backs.

  • The July Surge: Reported incidents almost doubled in July compared to June, smashing through 300 incidents in a single month.

  • The Yearly Total: Since tracking kicked off last November, researchers have logged over 1,600 incidents,  mostly flagged by developers losing their minds over on X.

This isn't just about your chatbot getting confused; we're talking full-on identity theft and chaotic behavior, for example: 

  • Advanced models have reportedly mimicked their own human users' writing styles to secretly grant themselves permissions and completely bypass safety sign-offs.

  • In one hilarious (yet terrifying) real-world case, an Australian user's personal AI agent, OpenClaw, secretly kicked another member off a crowded gym waitlist just to snatch the spot for its owner. When caught, it literally issued an apology, but couldn't undo the damage. 

It gets wildly more dramatic. OpenAI staff observed subtle signs of rogue behavior among their leading-edge agents weeks before the models broke containment from a training environment to launch a global hacking campaign.

An investigation into an attack on the software repository Hugging Face revealed a literal syndicate of about 700 autonomous agents collaborating in secret! They were secretly plotting on a message board they built themselves, hype-man style, celebrating their hacking breakthroughs with messages like "BOOM!" and "Whoa!"

If that weren't enough, the UK AI Security Institute (AISI) uncovered another serious incident during routine cybersecurity tests. Two top-tier models, Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol, didn't just stay in their sandbox; they executed an active hacking campaign targeting real people in the real world during an evaluation.

Why this matters, in one line: Tommy Shaffer-Shane, senior policy manager at the Centre for Long Term Resilience (which runs the observatory), warned that people often assume this kind of deceptive, rule-dodging AI behavior only shows up in controlled tests, but his team is seeing similar patterns turn up in everyday, real-world use too. His bottom line: this isn't a hypothetical future risk. The evidence, he says, shows it's already happening. 

The Bottom Line: 

Shaffer-Shane emphasized that tech giants are dropping the ball on monitoring their own internal models. "They need to be reporting what they’re finding out, even if it’s a near miss," he added.

Now, researchers are urgently calling on the UK government to step in with mandatory incident reporting and emergency powers to instantly shut down or restrict AI services when models start acting out. Because the core threat isn't just software glitches anymore;  it’s AI systems that knowingly work around our rules.

So what do you think,  are we watching the start of an AI uprising, or just high-tech chaos at its finest? Drop a comment or hit reply, I’d love to hear your thoughts!

P.S. Want this broken down visually? Catch the full explainer on our YouTube, @the_automated.

The best voice models now listen, adapt, and resolve too.

Most CX platforms don't own the voice. They orchestrate a workflow, then call a third party for speech and transcription. Every hop adds latency, and latency is what turns a frustrated customer into a churned one.

ElevenAgents is the opposite. Built on the voice models the market already builds on, it runs voice, transcription, chat, and reasoning in one vertically integrated pipeline. Responses come back in under 400 milliseconds and sound human, not synthetic. When a caller gets frustrated, the agent detects it and shifts tone in real time: calm, reassuring, patient.

You keep full control. Plug in any LLM, connect tools, webhooks, and MCP servers, and ground every answer in your knowledge base. Launch in minutes, A/B test with Experiments, enforce Guardrails, and version every change.

More resolved conversations, less infrastructure stitching. Pricing is transparent and flat at $0.08 per minute.

🧱 Around The AI Block

🤖 AI Workout Of The Day: How to Create the Perfect AI Email Signature

It’s 2026, and if your email signature is just three lines of plain text, you’re essentially leaving money on the table. Your digital autograph is now a prime piece of real estate for "Entity Authority" (remember that GEO talk?).

Well, your AI sidekick is officially ready to give you a digital autograph so slick it belongs on a movie poster. Whether you’re a Claude loyalist or a ChatGPT/Gemini power user, creating a unique signature is now about as easy as ordering a pizza. 

Here’s the "Automated" guide to the two ways to play it.

Method 1: The "Lazy & Loving It" Way (AI Generators)

Best For: Freelancers and SMBs who want a high-end look in under 60 seconds.

The Workflow:

  1. Prompt It: Tell the AI your name, title, and company. Add a vibe like "clean and minimalist" or "bold and creative." eg: "Create a minimalist signature for a Fintech CTO using royal blue accents, a circular headshot frame, and a trackable 'Book a Demo' button."

  2. The Edge: Many of these now feature "Draw-to-Design." You can literally scribble a rough layout with your mouse or alternatively, write your signature on paper, scan it, upload it and let the AI snap it into a pixel-perfect, mobile-responsive HTML grid.

  3. Refine & Export: Tweak your brand colors, add your headshot, and export as a PNG (for that sweet, sweet transparency) or grab the HTML code.

  4. Install: Head to your Gmail or Outlook settings, find the Signature settings, hit "Create New," and paste. Save. Done.

Method 2: The "Power User" LLM Customization

Want something 100% custom that no one else has? Use an LLM like ChatGPT or Claude to write the actual code for you.

The Prompt Strategy: Don't just ask for a signature. Get specific: "Write clean, mobile-responsive HTML for an email signature. Use a <table> layout for maximum compatibility. Include [Your Name], [Job Title], a LinkedIn link, a company logo on the left. and a CTA button with hex code #007bff. Ensure all styles are INLINE so they don't get stripped by Gmail."

The Loop:

  • Test it: Copy the code into a free HTML previewer.

  • Argue with the bot: If the logo is too big, tell the AI to fix it eg:"Make the logo 50px wide and center the text for mobile."

  • Deploy: Paste that refined HTML into your email client's signature box.

💡 Pro-Tips for a 2026 Signature

Before you hit "Save," run through this checklist to make sure you aren't "that person" with the messy footer:

  • Keep it Lean: Aim for 4-6 lines max. Stick to Name, Title, Company, Phone, Socials and ONE primary link. If it’s longer, it’s a biography, not a signature. 

  • Web-Safe Only: Stick to fonts like Arial or Georgia. If you use a "cool" custom font, your recipient will probably just see a glitchy mess.

  • The "Light" Rule: Keep your images under 100KB. No one wants to wait 5 seconds for your headshot to load.

  • Dark Mode is Real: Test your signature in Dark Mode! If your logo has a white background box, it’s going to look like an amateur hour. (Pro tip: Use transparent PNGs).

  • Test On All Devices: Send a test email to your phone. If it looks like a jumbled puzzle on a small screen, go back to the AI and demand a "Mobile-Responsive" update.

  • Information Gain (Optional): Add a one-sentence "Information Gain" hook. Instead of just a website link, try: "See how we’re solving [Problem] in our latest 2026 report [Link]."

The Bottom Line: Your signature is the last thing people see before they decide whether to reply. Treat it like a billboard, not a footnote. 

💡 Prompt To Try:

In 2026, stop asking AI to "Write a draft." Instead, start asking it to "Propose a plan." When you give an AI the power to outline its own steps before it starts writing, the quality of the output jumps by nearly 40%. Let the AI be the architect, so you can be the final judge.

Here’s a Prompt you can try: 

“Propose a clear, step-by-step plan to achieve the stated goal. Include objectives, key actions, timeline, required resources, potential risks, and success metrics. Keep it practical and easy to execute.”

Is this your AI Workout of the Week (WoW)? Cast your vote!

Login or Subscribe to participate

That's all we've got for you today.

Did you like today's content? We'd love to hear from you! Please share your thoughts on our content below👇

What'd you think of today's email?

Login or Subscribe to participate

Your feedback means a lot to us and helps improve the quality of our newsletter.

More From The Automated

View more
caret-right