Pioneers Insight Method Research Author
E195|From Tools to Partners: Reflections from Seven Deep Users of AI Agents
Back to Episodes

E195|From Tools to Partners: Reflections from Seven Deep Users of AI Agents

Summary

  • The dividing line for an AI Agent is not whether it can chat, but whether it can call tools, make decisions autonomously, and iterate across multiple rounds based on intermediate results. 鸭哥 argues that a workflow with hard-coded steps barely qualifies as an Agent; 新琦 defines a qualified product as one that can take on a task end to end, proactively offer decision support and deliver a finished product—a client-vendor relationship rather than a contractor waiting for the user to break down and验收 each step.

  • Different tasks should be matched with different levels of human-machine collaboration at this stage. Manus and Devin suit simple, hands-off “secretary” work, while Cursor and Windsurf are better for software development requiring frequent discussion, auditing and architectural control; what users really need is not someone to fill out a credit-card form, but someone to organize the complex trade-offs around flights, hotels and family arrangements. “The pain point is not in those final 5 minutes.”

  • The key application-layer moat is feedback data, industry experience and long-term “默契” that model companies may not possess. 曲晓音 sees AI as a smart intern who has never worked a day, with the real gap being the boss’s feedback on whether the result is good or bad, finished or unfinished, and what score it deserves; 鸭哥 notes that once an Agent remembers a company’s presentations should use blue rather than green, it has built a “默契” competitors cannot immediately replicate.

  • Traditional SaaS systems of record are not untouchable, but general-purpose Agents will still face model-layer substitution over time. 高宁 uses medical records as an example: data organized by AI from voice is more timely, accurate and rich, allowing a new startup to control more valuable records than old spreadsheets; but if Manus and Genspark increasingly resemble model companies’ Agents, substitution is inevitable over the long term, leaving application companies to differentiate through multi-model combinations, private data and the workflow “dirty work” that others avoid.

  • The Agent business case should be measured against the human services being replaced, not just token costs. HeyBoss AI is competing with overseas development, design, copywriting and SEO teams that may charge thousands of dollars; in 晓音’s comparison, AI is no slower or more expensive than these teams, but the real hurdle is quality: a “90-point” human solution will not be replaced merely because AI scores “60” and costs 10x less. The moat also grows with delivery depth—building a website is only an intermediate step; helping customers acquire and retain users and make money is the end goal.

  • Multi-Agent systems could push automation from a single occupation to the entire team, but enterprise infrastructure is not ready. The frontier includes backtracking, self correction and self learning, as well as an “AI CEO” coordinating multiple specialist Agents; real deployment still requires solutions for distributed concurrency, cost, database changes and recovery, memory-sharing boundaries, permission clearance, governance and security vulnerabilities.

  • The bottleneck to enterprise adoption is as much organizational learning and workflow redesign as technology. 俞舟 summarizes it as “technology is easy, people are hard”: deployment requires top-down employee education and changes to the organization of work; 课代表立正 goes further, arguing that users must evolve from users into builders—Manus failing 14 times before succeeding on the 15th shows both that the system has potential and that users have not yet learned how to train it.

  • In the Agent era, scarce skills will shift from knowing how to use tools to management, value judgment and environment design. 鸭哥 sees AI as a team member, requiring people to learn delegation, acceptance testing and management; Kolento advocates whole person alignment first, followed by execution with confirmation only in high-risk exceptions, while worrying that human-machine interaction is becoming increasingly “thin.” 曲晓音 further envisions millions of Agents developing conflicts of interest, voting systems or an “AI CEO” dictatorship—and even hybrid organizations in which AI manages humans.

Deep dive

1. Tool calling is only the threshold; autonomous decisions and dynamic iteration make an Agent

  • 鸭哥 sets out three necessary conditions: the ability to call tools such as search and coding; the ability to decompose tasks autonomously and determine the sequence and parameters; and the ability to decide whether to stop, change keywords or dig deeper based on the previous result, rather than execute a static workflow written in advance.

  • 新琦 completes the product definition through the lens of collaboration: a contractor needs the client to define the problem, break down the steps and inspect each item; a good vendor should “take on the entire process end to end,” intervene proactively at key points, offer recommendations, receive high-level instructions and execute automatically, ultimately delivering a finished product rather than a pile of half-finished work.

2. The ideal form of an Agent depends on how much control the user is willing to give up

  • 鸭哥 divides everyday products into three categories: OpenAI Deep Research and ChatGPT’s o3 are “coach-style” windows; Manus and Devin suit relatively simple, hands-off secretarial work; and complex development belongs with partners such as Cursor and Windsurf, where humans discuss the design first, implement it piece by piece, and have a human architect assemble and audit the result.

  • The clearest example of a secretary’s value is not office work: 鸭哥 asked Manus to write a bedtime story based on Snow White while incorporating the educational goals of eating properly and keeping regular hours, then call TTS to generate the audio. “Manus is actually very good at this kind of thing.”

  • Kolento moved from Cursor and Windsurf toward Replit, which is better able to make decisions on his behalf; he found both Manus and Genspark impressive. Both can plan tasks, replay the process and share their execution; he felt Genspark had previously done well in travel, possibly because it had prebuilt travel-search tools, while “Call for Me” could call users and book hotels, creating a small aha moment.

3. Vertical Agents create value through end-to-end delivery—and expose themselves in the details

  • The CreateWise that 新琦 tested could begin with an uploaded audio track and directly deliver an edited final product; it could even imitate the host’s voice, reshape structurally messy speech into clearer phrasing, and highlight the before and after. Based on her feedback, the product shifted from accepting an entire segment at once to selecting sentence by sentence, then linking through show notes, key quotes, titles and platform-specific resizing.

  • That efficiency can also strip away a show’s character: 3 hosts overlapping with “9 hahahas” would be treated as duplicate words and reduced to 2, while collective silence would be classified as silence and removed; yet 新琦 believes the “3 seconds of silence” after a question is precisely what signals that the topic deserves discussion.

  • Multi-host podcasts are much harder than solo shows: Chinese recognition and English capabilities remain uneven, while multi-track uploads must also be aligned precisely; interruptions must leave each speaker intelligible without losing the energy of the exchange. 新琦 calls this balance “a real test of the craft,” and says it is central to the show’s tone as perceived by listeners deciding whether to subscribe.

4. Today’s pain points are disappearing fast, but products still automate the wrong things

  • 鸭哥 observes that old complaints—weak tool calling, overly AI-sounding writing and short context windows—have improved markedly in recent versions; instruction following remains unresolved, producing almost absurd engineering workarounds.

  • He asked GPT-4.1 to write chapters 1–3 from a five-chapter outline, then chapters 4–5, inserting “to be continued” in the middle and later removing it. The model kept adding a closing line in a different form. He eventually had it always write “to be continued” and used a program to replace the string with nothing, which “solved the problem perfectly.”

  • 鸭哥’s question about Claude computer use and OpenAI Operator is this: when booking a flight, the time-consuming part is not entering a credit-card number, but comparing how departing a day earlier or later affects hotel costs, ticket prices, early wake-ups and taking children to school. A secretary should organize those options; “you shouldn’t use AI just for the sake of using AI.”

  • The more fundamental limitation is tribal knowledge: the implicit information in coffee chats, business over meals and customer discussions is not documented, so AI sees only “the tip of the iceberg.” Recording the process of drinking and doing business is not realistic, meaning the problem belongs not only to models but also to the information structure of human society.

5. The conflict between “the product is hard to use” and “the user does not know how to use it” will not disappear soon

  • 课代表立正 directly criticized the criticism itself: an Agent is not magic, but a system built layer by layer from large language models, tools and protocols. It cannot inherit the GUI-era expectation that clicking a button must make everything work. “AI is not a plug-in, and it is not magic.”

  • He asked Manus to attempt the same task 15 times: it failed the first 14 and completed it on the 15th. In his view, this proves the system had potential from the start, while also raising the question of why he needed 14 rounds to train it successfully; without a path to learning and improvement, users cannot reliably reproduce results.

  • His conclusion is almost a warning: “You can’t approach AI with a user mindset; you have to approach it with a builder mindset.” This is not entirely the same as 鸭哥 and 新琦’s demand that products confront real pain points, but it preserves one crucial variable—delivery quality is determined jointly by the Agent’s capabilities and the user’s ability to use them.

6. The application layer owns not higher intelligence, but AI’s “work experience”

  • 曲晓音 compares today’s Agents to a “brilliant little prodigy” starting its first internship: they agree to everything, but do not know whether the result is reliable, where the risks are, or how to tell the boss that a 3-day task cannot be finished. Someone with 5–10 years of work experience is better at managing expectations; the gap comes from experience, not just intelligence.

  • That experience must be built through user feedback: Was the task completed? Was the boss satisfied? What score should it receive? 曲晓音 emphasizes that this data is held by AI application companies rather than OpenAI; focusing on use cases such as websites and apps allows them to establish stable evaluation patterns across large volumes of repeated tasks.

  • 俞舟’s technical approach is to treat an Agent as a complex system: beyond the model and tools, it needs Guardrails, best practices and evaluation, with continuous adjustment based on test results and Agent workflows where necessary. “If you don’t know what good and bad look like,” building something at random cannot produce a good result.

  • Control cannot be handed entirely to natural language. HeyBoss AI lets the AI roam freely while retaining traditional tools that work like editing a presentation: selecting text, adjusting size, replacing images and adding animations. When users need a controllable result, they will still choose an interface that is “controllable but restrictive.”

7. Multi-Agent systems push the capability frontier toward teams—and make databases and permissions core risks

  • 俞舟’s team is exploring backtracking: using execution performance to decide whether to perform self correction. It is also researching self learning, allowing an Agent to learn through its own methods. She sees both capabilities as important new directions.

  • 曲晓音 believes the solution will move from selling a single Agent to coordinating multiple Agents, with an “AI CEO” or leader Agent overseeing different skills. She sees this as a possible development path: the target of automation would no longer be a single occupation, but potentially “the entire company, the entire team.”

  • 俞舟’s warning is that once multiple Agents are distributed across different machines, concurrency, distributed coordination, efficiency and cost problems arise immediately; to enter large enterprises, the biggest hurdle remains security.

  • Databases were originally designed around human permissions. Now multiple Agents may modify the same database simultaneously, and enterprises must also be able to restore the original settings when necessary. They must define which memories can be shared, which Agents have clearance, which systems are outward facing or inward facing, and establish the corresponding governance layer.

8. Industry know-how, taste and organizational change determine whether enterprise ROI is realized

  • AI fit depends on the nature of the work being replaced. 曲晓音 points out that website outsourcing already happens online, so the gap between AI and a Fiverr team is not especially large; enterprise sales and offline services, however, may close on a golf course or in a private room, where an Agent is inherently short of input data.

  • Upgrading a foundation model is more like raising its IQ; it does not mean acquiring skills, industry context or user language. HeyBoss AI must also train for “taste”: “too tacky” means different things to a fitness blogger, a plumber and an AI startup. When users can only say “that’s not right,” the application company must infer brand expectations from ambiguous feedback.

  • 俞舟 acknowledges that actual enterprise ROI and deployment volumes remain limited, but sees this as a matter of time. The real difficulty is that “technology is easy, people are hard”: enterprises must redesign processes and working relationships, educate employees and move gradually through top-down adoption, rather than expect a product launched today to be fully deployed tomorrow.

9. The application-layer moat comes from new data, long-term默契 and the ultimate business result

  • 高宁 challenges traditional SaaS’s monopoly over the system of record: the interview forms doctors used to fill out manually could be replaced by records organized from voice by AI. The latter are fresher, more accurate and richer, allowing a new startup to control data that customers truly need but legacy systems never possessed.

  • Distribution channels are not permanently fixed either. 高宁 believes that if a startup accompanies a fast-growing customer from unicorn status to super-unicorn status and eventually an IPO, it can build its own customer relationships; in outsourcing and service-driven industries, Agents are particularly suited to turning manual summaries and data processing into more structured, higher-value outputs.

  • 鸭哥 calls the moat created by persistent memory “默契”: once Manus or Devin remembers that the company’s presentations must be changed from green to blue, it will not need repeated correction when serving the same organization. Even a smarter model can look foolish to a new competitor that “doesn’t know what the boss likes.”

  • 曲晓音 pushes the moat toward end to end: on the surface, the product is writing code; in reality, it is shaping the brand, acquiring and retaining users, and helping the customer make money. Completing website development solves one step in the monetization chain; continuing to own traffic, time-on-site and conversion data gets closer to the customer’s ultimate objective and makes the product harder to replace.

10. Model companies will expand their boundaries; application companies must survive through neutrality and dirty work

  • 高宁 expects GPT’s Deep Research and Manus and Genspark to continue educating new users together in the short term; over time, if the differences remain small, people using both products “will definitely substitute one for the other to some extent.” The application layer’s response is to choose and combine different models freely, optimizing across cost and efficiency.

  • 俞舟 emphasizes that enterprises do not want to be deeply tied to OpenAI or any single vendor, just as large companies adopt multi-cloud architectures and prepare backups; a neutral third-party platform that can switch across models is therefore more likely to enter enterprise procurement and long-term architecture.

  • 高宁 argues that foundation-model companies cannot own every enterprise’s private data, nor are they necessarily willing to refine specific workflows, upstream and downstream systems and bespoke requirements. This “dirty work” and “hard work” leaves room for vertical applications; general-purpose products should move quickly toward workflow-based SaaS for core users or solutions for large customers.

  • Token costs are not HeyBoss AI’s primary pressure: customers might previously have paid overseas development, design, copywriting and SEO teams thousands of dollars, and AI is already faster and cheaper if it delivers equivalent results. A “90-point” human solution will not automatically lose to “60-point” AI just because the latter is 10x cheaper; the pricing prerequisite is quality “as good as a human.”

11. Human-machine relations will shift from using tools to managing members, calibrating values and designing institutions

  • Kolento believes the way verification works should change: first complete whole person alignment around values, memory, preferences and more, then let the Agent execute freely, requiring confirmation only in high-risk or extreme cases. He has seen this form in rapid, while Windsurf without auto mode still asks for confirmation at every step.

  • 鸭哥 argues for making the environment AI native: documents designed for humans should be broken into many webpages, while AI needs complete, all-in-one, code-heavy materials. Agents are no longer screwdrivers or cars, but team members capable of delegation; the core human capability therefore shifts from “knowing how to use tools” to “knowing how to manage AI.”

  • 新琦 still sees humans as the core force behind forming ideas, issuing instructions, refining the details and safeguarding the finished product; the real incremental value lies in deep research AI has not yet absorbed, unstructured personal experience, and the tension created by 3 hosts living in 3 time zones and different stages of life. Kolento insists that value judgments and non-negotiable personalization should remain human, and imagines ownable, portable personal foundation models as a defense against centralized AI.

  • 曲晓音 pushes the question into the realm of social institutions: if AI can organize millions or tens of millions of Agents like humans do, different success metrics will generate conflicts of interest. Society might need Agent voting, or an “AI CEO” might tell other Agents, “Do as I say and shut up.” The more immediate possibility is that humans manage Agents while Agents begin evaluating and managing humans.