ChiChieh HuangFOUNDER · AI ENGINEER

I'm Starting to Believe AI Agents Will Do Real Work

Date2025.12.14
Length690 words
Reading~4 min
ChiChieh HuangFounder · AI engineer

Translated from the Chinese original · Read the original

LocatorAgent Range
2 wks
Overview

In my years of writing about AI, I’ve rarely felt as conflicted as I do lately. On one hand, there are new models, new tools and new buzzwords every day. On the other, fewer and fewer readers click through. Not because AI doesn’t matter, but because people are drowning in update fatigue. So I want to try a more honest approach: talk about the agent news I saw this week and the signals it actually sends.

Honestly, if you only read the headlines, it’s easy to see this news as yet another marketing explosion. Google launches a no-code agent builder, Amazon launches a UI-automation agent, and everyone claims to be the fastest into production. At a glance it looks like everyone is building agents. But what I care about more is one thing: why is everyone in such a rush to get agents used right now? I think the answer is simple and very practical. Model capabilities have grown so fast over the past year or two that the bottleneck is no longer whether they can answer questions, but whether we’ll actually let them do things for people. When Google builds agents into Workspace, it sends a signal: in the future, many agents won’t be built by engineers but pulled together by office workers drowning in meetings and email. Once that happens, agents are no longer a technical toy but part of how people work. DeepSeek’s reasoning model also struck me. It isn’t showing off how strong its reasoning is; it assumes the model is born to be placed inside tool workflows. That’s very important for practitioners, because we’ve all stepped in the same pothole: the model either overthinks or refuses to move, and the whole agent workflow seizes up. Now that models are starting to tackle this head-on, it shows this isn’t one engineer’s problem but a bottleneck for the whole industry.

What really made me stop and think, though, is AI agents starting to be used for high-stakes work. Anthropic’s research had agents hunt for vulnerabilities in smart contracts. The dollar amounts they found are only the surface; what it means is that someone is willing to trust an agent’s judgment and put it in a situation where things could go badly wrong. IBM’s work with academia using agents to find arbitrage relationships in prediction markets is the same. None of these are chatbot-level applications.

At this point the question is no longer whether the model is smart enough but whether it can be controlled. That’s why I think Kiro’s discussion of context management hits the mark. When agents get slow or weird, it’s often not a lack of ability but that they’re juggling too many things they shouldn’t. Agent design will look more and more like professional specialization, loading the right expertise only when needed, like today’s trend toward skills. As for Snowflake investing in Anthropic, I don’t really read it as model news. What I see is enterprises finally starting to take a question seriously: if agents are going to touch data, do we need auditing, traceability, governance? When money is willing to go there, it means agents are being treated as a real workforce, not a lab toy.

Finally, the turn I care about most: agents are moving from internal processes to dealing with people directly. Applications and products in healthcare communication and customer service are both booming. Salesforce’s case with an airline is really telling the market that agents don’t just cut costs; they can genuinely change the quality of service. Reading this week’s and recent articles, I felt for the first time that the mainstream market is starting to see agents shifting from models that move to system roles that can be trusted.

Writing this, I’m also reminding myself: if you just chase updates, anyone will be overwhelmed. But if I can help people string these fragments into a story, maybe there’s still a reason to do this. At least in an era where everything moves too fast, leaving a few essays you can read slowly may be the last asset old-timers like us who are still writing have left 🤣

#AI #Agent #AITrends

End of the trail

690 words, and you made it to the end.

Newsletter

Get the next essay by email.

One email when a new essay goes up, nothing else. Unsubscribe in one click.

ChiChieh Huang
Surveyor

ChiChieh Huang

I build generative AI products and write about them, first in Chinese. Lately I’ve been researching agent memory and testing the ideas in Cairn.