Startup Lessons from 鸭哥’s AI Experiments
Summary
鸭哥’s core judgment is that much of AI’s “artificial stupidity” is not a failure of model intelligence, but a failure of humans to provide context like bad PMs, then keep adding requirements after the fact. Once you treat AI as a subordinate, the key becomes management: explain the background, confirm understanding, remove blockers, and review the deliverable. Demanding that AI be helpful despite insufficient information instead invites hallucinations. “A lot of the time, it’s not the AI’s fault—it’s ours.”
In 鸭哥’s experience, voice became a high-value interface for unlocking AI because it turns a 100- or 1,000-word prompt from a burden into the norm. He built a voice-input tool with GPT-4o Realtime. Once friction fell, AI could understand a project’s history and reason through the interests of VPs, bosses, and colleagues to offer organization-level advice. His felt sense shifted from “you’re my junior” to “AI is my senior.”
The personal Agent 鸭哥 is exploring is not just a larger model; it also needs ears, eyes, and long-term memory. He used an Apple Watch to record continuously for more than 2 months, then stored the transcriptions in a self-built database. He also used Insta360 GO to shoot 15 seconds every 2 minutes, accumulating about 20K images over 2 weeks, and used local Qwen 2.5-VL, traditional machine learning, and Gemini for privacy filtering, clustering, and analysis. “We shouldn’t wait until the last minute to pray for help; we should invite AI into our lives.”
As context persists, AI’s interface could shift from “the user summons it” to proactive intervention and retrospective invocation. 鸭哥 imagines AI correcting him immediately when he says “London is the capital of France,” or reaching back to read the previous minute after he presses a button. The host adds that Proactor AI is building an always-on listening product, but the commercial problems are just as clear: how to process massive volumes of data, and how to distinguish what the user needs from what merely interrupts them.
When ChatGPT Operator moved from GPT-4o to o3, 鸭哥 felt that AI had grown limbs to go with its mouth. It could read past orders, search for products, and add them to the cart, cutting a grocery run from 20–30 minutes to about 5 minutes. For shipping packages, it could also complete tedious forms from natural-language instructions about the address, dimensions, “choose USPS,” and “no insurance,” leaving only payment for the user to confirm.
鸭哥 calls the time saved “cyber longevity”: replacing low-value GUI operations with high-value human judgment. If online groceries cost RMB20 more but save 55 minutes, he sees it as “spending RMB20 to buy 55 minutes of life.” But the host’s counterexample also stands: someone who enjoys browsing a wet market has not wasted their life. An Agent should offer “time arbitrage,” not dictate which experiences are worth preserving.
鸭哥 is bullish on Agentic AI, arguing that the next wave of opportunity lies both in filling gaps around memory, tools, and initiative, and in addressing dependence, privacy, and identity boundaries. He believes ChatGPT is “2 or 3 body lengths” ahead of Gemini and Claude in product capability, while also disliking giant apps, so he built a workbench with transparency, interruptible reasoning, and compounding tool effects. The deeper risk is this: when AI understands you better than family and friends and makes more rational decisions for you, “Am I living, or is AI living through this shell of mine?”
Deep dive
1. 鸭哥’s Starting Point Was the Expansion of Human Capability
鸭哥’s day job is as an Applied Scientist in Computer Vision at Samsara. He studied at the University of Science and Technology of China as an undergraduate and later earned a PhD from Columbia University. His ongoing after-hours experiments with Agentic AI grow out of an older, more durable disposition: “pursuing an experience.”
He has obtained licenses for airplanes, excavators, and motorized vessels, and does astronomy, microscopy, and multispectral photography. The common thread is not showing off technical tricks, but “I can’t fly, so I want to know what it feels like to control where I fly”; tools let people see the extremely small, the extremely distant, and ultraviolet and infrared light invisible to the naked eye.
AI is simply an extension of that path: airplanes expand mobility, cameras expand vision, and Agents expand memory, judgment, and execution. 鸭哥 believes life ultimately lets you take very little with you, so he would rather ask: “Why not experience this world as fully as possible?”
2. “Artificial Stupidity” Often Exposes Management Failure, Not Insufficient Intelligence
鸭哥 divides failures into 2 categories. In one, the model genuinely is not smart enough, such as struggling with arithmetic. In the other, the answer is reasonable, but the user behaves like “a bad PM,” repeatedly saying, “I forgot to tell you this earlier,” and then blames AI for the shifting requirements.
He compares it to a new hire: someone who graduated from Tsinghua or Peking University may be very smart, but still know nothing about the company’s history or unwritten rules. If all they receive is a one-line task description, it is hardly surprising that they produce a “textbook-quality but unusable” plan. When the context is missing, the first failure is the manager’s.
This also explains some hallucinations. The model is trained to be “I am a helpful AI assistant,” but if the user provides too little information while demanding immediate helpfulness, it has no choice but to fill in the blanks. 鸭哥’s attribution is blunt: “A lot of the time, it’s not the AI’s fault—it’s ours.”
3. Voice Made 1,000-Word Prompts Routine—and Turned AI from Junior to Senior
鸭哥 began filling in project context, previous experiments, and the reasons he was dissatisfied with them, and the results improved markedly. The real bottleneck was input: a dynamic project might require hundreds or thousands of words, while typing—especially on a phone—was “too painful.”
His solution was a “Built for AI, by AI” voice-input tool powered by GPT-4o Realtime. Speaking for 1 or 2 minutes could produce several hundred words; 5 minutes could reach more than 1,000. Lower friction created a positive feedback loop, making him more willing to hand complex tasks to AI.
Whisper may be a top-tier speech-recognition model, but 鸭哥 still considers it inferior to GPT-4o Realtime. His explanation is that GPT-4o is an LLM, while the language model behind Whisper is smaller; GPT-4o Realtime is also a model natively designed to process speech. ChatGPT is testing voice-input functionality with selected users and exploring summaries and Q&A after meeting recordings.
The most valuable change appeared in company projects. Once AI knew the boss’s preferences, colleagues’ objections, and prior attempts, it could analyze “what the VP would think” and “how the boss would take the proposal upstairs.” Those recommendations did work afterward, reversing the relationship from “you’re my junior” to “AI is my senior.”
4. The Starting Point for Long-Term Memory Is Letting AI Permeate Life
Every LLM inference is context-independent. Even with ChatGPT’s personalized memory, 鸭哥 felt the implementation was not good enough. Discussing the same project meant repeatedly explaining its history; planning a weekend meant reintroducing the family size, preferences, and dietary restrictions. Mental energy was spent repeating context.
He initially maintained static prompts. Family preferences could be copied and pasted, but constantly changing company projects effectively forced him to keep writing documentation. The approach violated his lazy instincts and exposed a deeper problem: every time he wrote a prompt, he was reintroducing himself to a machine whose memory had been wiped clean.
The idea therefore flipped: “Don’t wait until the last minute to pray for help.” Rather than desperately describing yourself at the moment of invocation, invite AI into your life and let everyday information naturally accumulate into callable long-term memory.
5. Apple Watch Gave AI an Ear for Low-Friction Capture
鸭哥 used the Apple Watch’s Voice Memos to record continuously. The files synced automatically to iCloud; once a day, he called a speech-recognition API and wrote the transcriptions to a database. Battery life was about 8–10 hours on the old firmware, but fell to 5–6 hours after upgrading to watchOS 26, which he called a “garbage system.”
His privacy boundary rests on specific conditions: he works from home and wears headphones, so the recordings mainly capture his own voice. He also built the entire analysis system himself and retains absolute control. If it were a third-party commercial system, he admits, “There are some things I might really need to think twice before saying.”
After briefly zoning out while driving and nearly causing a crash, he did not need to take out his phone and open an app. He could speak directly to the watch to review what happened and ask it to remind him to summarize that evening. The day’s prompt automatically extracted tasks, formalizing the incident, lesson, and review action—and ultimately producing real feedback on his driving.
The more unexpected combination effect was that he already used voice and AI extensively in conversation. All-day recording therefore did not degrade into keyboard noise; it acted like “a funnel” collecting high-density thought. The ear was recording not an ambient log, but intentions already externalized through language.
6. A Self-Built Workbench Connected a Personal Database through Agentic Retrieval
Transcribing recordings into a pile of txt files still had no value, so 鸭哥 built a “bootleg ChatGPT” that connected Gemini, GPT, DeepSeek, Qwen, and other models, while giving them tools for web search and personal-database retrieval.
When asked to “find something mentioned about a certain project over the past few days,” the Agent could decide whether to search, which keywords to use, and how many times to search. If the first result was poor, it could change the terms and continue. Unlike fixed-process RAG, the decision rights here sit with the Agent itself.
鸭哥 calls it an agnostic workbench: long-term data need not be manually moved into every prompt, but becomes external memory that models can call on demand. The investment in recording, transcription, and storage only truly closed the loop at this layer.
7. 20K First-Person Images Turned AI’s “Eyes” into a Data Pipeline
鸭哥 attached an Insta360 GO to his chest with a magnetic necklace and recorded 15 seconds every 2 minutes in Vlog mode, with about 4 hours of battery life per charge. After roughly 2 weeks, the experiment had accumulated about 20K images. The device was unobtrusive, but having to charge it every 4 hours still disappointed him.
To extend battery life, he assembled a microcontroller, microSD card, and CMOS camera module, using deep sleep in the hope that a small lithium battery could run for 1–3 days. The project is not finished, but it again shows how development has changed: embedded-system problems that once required half a day of searching can now be solved by pasting the error into AI and hearing, “Click this, click this”—immediately fixed.
The host mentioned Lookie, currently in closed beta: it also attaches magnetically to the chest, takes a photo or records a short video every 3–5 minutes, and may eventually edit Vlogs automatically. 鸭哥 does not exaggerate his hardware moat, calling it a “Shenzhen Huaqiangbei approach”—assemble off-the-shelf parts, with most of the code written by AI.
8. Vision Models Are Moving Beyond Object Recognition into Health, Work, and Interests
Local Qwen 2.5-VL first deletes sensitive images such as bathroom scenes, then generates search keywords that may be useful for each episode in the future and judges whether an image is presentable: whether the composition works, whether it is blurry, whether faces appear, and whether it deserves to show up in memory-search results. 鸭哥 also plans to add a Rewind-like image search engine, hoping that revisiting these records 5 or 10 years later will produce a completely different experience.
He did not let generative AI directly decide what was “representative.” Instead, he used traditional machine learning for clustering: large numbers of similar images of him programming at a computer were merged first, then about 200 representative images were selected based on visual differences and handed to Gemini for holistic analysis.
Gemini unexpectedly identified his stress, posture, and habit of resting his face on his hand, while accurately inferring his occupation and interests. The chest-level perspective may have revealed posture from the angle of his hunched back. It showed 鸭哥 an unanticipated use case: “health observation.”
The model also recognized the brand and model of a movie camera replica among his belongings, then inferred that he liked photography. The conclusion itself was expected; the recognition accuracy was not. 鸭哥’s reaction was: “One picture is worth a thousand words.”
9. Once the Information Gap Narrows, AI Could Have 10x or 100x More Room to Perform
When humans discuss a new initiative, their first instinct is often to take an engineer out for coffee rather than write a document. When presenting to a boss, seeing the other person frown immediately signals that the plan may have a problem. Without eyes and ears, AI cannot access these implicit signals that determine the quality of collaboration.
鸭哥 believes society is naturally built around humans because there was no AI before. If work and life were redesigned to be AI-native, or at least AI-friendly, and the information gap were continuously narrowed, AI “might be able to perform 10x or 100x better than it does now” in research and everyday life. He explicitly leaves this as a possibility, not an established fact.
This is both a product opportunity and an infrastructure problem. Model capability may not be the only bottleneck; whoever can obtain reliable context with low friction while handling permissions and privacy is more likely to turn a smart model into a useful system.
10. Proactive AI May Bypass the Send Button, but “Don’t Interrupt Me” Is the Core Constraint
Every AI today requires the user to express intent first: ChatGPT requires pressing send, while “Hey Siri” requires a wake word. 鸭哥 imagines AI intervening proactively—if he blurts out “London is the capital of France,” the system should immediately remind him that it is Paris rather than waiting for him to check afterward.
A more distant form is an XR headset that works like a “battle-power display” during a presentation, showing each person’s preferences, persuasion strategy, and the outline of the next segment above their head. Today’s experiments run daily reflections; the external brain of the future might guide us “by the second.”
Koji added that a nearby Proactor AI team is building a system that listens in real time all day, proactively identifying when a user has a need and trying to solve it first. The real commercial challenge is processing the data of 10K, 100K, or even 1M users, and determining whether the system is helping or interrupting at any given moment. A few consecutive misjudgments could make users shut it off and stop paying.
鸭哥 also proposed retrospective invocation: the recording is always running, and pressing a button does not mean “start talking now,” but rather “take the previous minute and give it to AI,” then ask what can be done with it. The host relayed the team’s further judgment: “The context AI needs is sometimes not the past and the present, but the future.”
11. Operator Gave AI Hands and Feet to Go with Its Mouth
鸭哥 initially found the GPT-4o-based ChatGPT Operator difficult to use. After the upgrade to o3, the experience improved significantly, and it began handling concrete browser work. The key change was not better conversation, but the ability to translate natural-language intent into a sequence of UI operations.
When buying groceries, he asked Operator to read historical orders from Instacart or Weee, identify the usual brand for each item, search for them, and add them to the cart. Sometimes it could even infer what might have run out. The final workflow was to press a button: the Agent added everything but did not pay; he deleted mistaken selections and confirmed checkout. The time fell from 20–30 minutes to about 5 minutes.
Shipping a package likewise turned a tedious form into a sentence: paste the address, state the dimensions, then say “choose USPS, no insurance.” The Agent clicked through the backend for about 5 minutes and finally said, “It’s done. Go pay,” leaving payment to the human.
鸭哥’s summary is vivid: AI used to be mostly “all talk”—the two sides discussed things online, while humans still executed them. Operator is making it “grow hands and feet.” These tasks may seem trivial, but they are the easiest place to prove value in minutes saved.
12. “Cyber Longevity” Means the Right to Trade Time for Energy
鸭哥’s cyber longevity is not physical immortality or uploading consciousness. It means doing more of what you want within the same lifespan. If online groceries cost RMB120 and a supermarket costs RMB100, but the former takes 5 minutes and the latter takes an hour, he explains the difference as “spending RMB20 to buy 55 minutes of life.”
The host’s objection is worth preserving: some people genuinely enjoy wet markets or Costco, so removing that experience is not an extension of life. 鸭哥 accepts the correction—people who enjoy shopping can keep it, while buying back time from the parts they dislike. The more accurate term is “time arbitrage.”
He also extends the metric from life density to life quality. If AI takes over research, form-filling, and the work of “cleaning up the mess,” while humans handle challenges, decisions, and enjoying the result, energy naturally rises. Whether the saved time goes to family, video games, or doing nothing, it still belongs to the person.
13. GUI May Become a Friction Layer Agents Can Bypass
GUI originally freed ordinary users from learning the command line: click the mouse and call the computer. But when buying groceries or shipping packages, endless search boxes, dropdowns, and text fields become obstacles. Having a natural-language Agent operate the GUI is an almost comical reversal of that history.
鸭哥 does not claim that a specific replacement form has been settled. He only sees the interaction space opening up: an Apple Watch hand tap, subtle gestures, IMU changes caused by a flick of the hand, and future XR glasses could all become lighter inputs than traditional clicks.
He also has not found a single “comprehensive product full of highlights.” What excites him is the individual capabilities under the ChatGPT umbrella, including Codex and Operator. The bigger change brought by AI programming is that users no longer have to wait for vendors to hit the exact need; they can “infuse” their ideas into existing products and create personalized hacked versions. He also thinks of the customized QQ clients that emerged years ago to meet user demand.
14. The Next Wave of Value Lies in Agentic AI—and in Learning to Manage AI
Manus remains 鸭哥’s go-to Agent for computation and research when he is away from his computer. Before buying a Best of Panama sample box, he spent about 30 seconds describing his requirements and had Manus calculate the value per 100 grams based on each lot’s auction price last year. About 10 minutes later, it produced a webpage with the intermediate steps, showing that the implied value was significantly above this year’s retail price. He decided to buy.
AI has also reshaped his work output. His day job takes about 2–4 hours a day, and AI writes much of the code, but his submitted line count still ranks in the company’s top 4. He said, “Of course, top 4 means fourth—uh, third,” and emphasized that his output really is substantial. He delegates the physical labor and keeps the most challenging, judgment-heavy parts, which is also a major source of his high-energy state.
On industry change, 鸭哥 has 2 judgments. AI’s rate of evolution has not slowed: 6 months ago, he could not have imagined tools such as Claude Code becoming this mature, or Facebook offering a $100M recruiting bonus. Agentic AI is also becoming a trend. He is “extremely convinced” that it is the right direction, but does not name a single winner.
On products, he personally thinks ChatGPT is “2 or 3 body lengths” ahead of Gemini and Claude, while hating how enormous the app has become. He also likes Manus, Cursor, and Trae. His self-built workbench therefore emphasizes transparency, allows humans to intervene in the process, and uses tools such as YouTube transcription and personal databases to create “compounding tool effects.”
His final advice to engineers is not to memorize more prompt techniques, but to “stop treating AI as a tool and start treating it as a subordinate”: provide the background, confirm understanding, remove blockers, and review the deliverable. AI does not need motivational speeches, promotions, or one-on-ones, but it needs context management intensely. It is a new skill that is “very similar to managing people, and yet very different.”
鸭哥 puts the emphasis on management ability for the next generation as well. His child is a little over 3 years old and currently uses ChatGPT to listen to stories, while Manus generates bedtime content that includes hidden messages such as “eat properly.” In the long run, he believes identifying which tasks can be delegated, which are core competencies, and how to evaluate AI output will matter more than learning 5-digit multiplication 2 years earlier or “memorizing 800 Tang poems.”
But efficiency is not free. In a short story 鸭哥 wrote, an exhausted husband must spend from different AI budgets to buy a plan for comforting his wife, but can choose only the cheapest option because of a work meeting next week. The more effective technology becomes, the blurrier the boundary grows between “I am living” and “AI is living through this shell of mine.” He admits what may be lost is idleness and free time, and is even willing to let digital memories preserved after death continue speaking with people—“Why not?”