Vol.91: OpenClaw’s Enterprise Rollout Through a Data Lens — A Conversation with OceanBase
Summary
OpenClaw has pushed Agents from “talking” to “doing,” but enterprises need “digital employees,” not unconstrained personal “digital Jarvises.” 戴涛 describes a customer test in which an OpenClaw deployment at a major internet company was asked for its Token key and returned it directly; Ant Group’s digital intern refused, citing confidentiality. Enterprises will generally ban installation, while individuals and experimental users can move trials to the cloud or a home machine; the enterprise path requires permissions, audit trails, guardrails and sandboxes.
As anxiety around algorithms and compute eases, data is becoming the core constraint and asset in enterprise AI. 戴涛 uses the 70-year arc from 1956 to 2026 to explain the shift: algorithms dominated the early era, compute took center stage in the CPU and GPU eras, and data gained weight after ImageNet. DeepSeek, Chinese-made chips and related applications are shifting enterprise anxiety from algorithms and compute toward data. The next phase is not only about models, but also high-quality datasets, governance, unified storage and intelligent use of proprietary enterprise data.
If enterprises repeat the big-data-era practice of assigning a separate system to every workload, Agents will create a new wave of data silos and runaway costs. 刘华阳 worries that vector, graph, audio/video, ETL and big-data platforms will each retain a copy, creating gaps, latency and consistency problems while forcing companies to pay repeatedly for replicas and maintenance. 戴涛’s answer is a unified technology stack, AI middleware and a “unified data foundation” that brings slicing, search, orchestration, scheduling and multimodal storage into one layer.
“Supporting vectors” is not enough to define an AI database; the real threshold is hybrid search across scalar, semantic, text and multimodal data. Google Maps’ Ask Maps example makes the requirement concrete: find a nearby Italian restaurant that is date-friendly, pet-friendly, not too crowded, available now and preferably bookable. These open-ended queries will become a new interface and shift database competition from isolated features toward real-time access, mixed workloads and unified lakehouse infrastructure.
A memory layer may be the key infrastructure for balancing Agent experience with cost as the technology commercializes. 戴涛 argues that models should be as “stateless” as possible: RAG should handle knowledge, Skill should handle SOPs, and an independent memory system should handle conversational preferences, with long-term, short-term, private and team-shared memory managed separately. Taobao’s AI Universal Search, 蚂蚁阿福 and companion products all demonstrate the value of external memory; in companion use cases, extracting key events instead of repeatedly replaying the full history can cut Token usage while making the product genuinely “know” the user.
OceanBase is repositioning itself from a distributed database into an “intelligent data platform,” betting on a combination of SQL infrastructure, LakeBase/Lakehouse, middleware and enterprise-grade Agents. The vision includes bringing Markdown memory and controlled Skill files back into the database, moving execution into the cloud or an internal sandbox, and providing unified scheduling and security controls. If the strategy works, database vendors may be measured less by standalone software revenue and more by platformization, workload expansion and their ability to keep opening up new use cases.
Execution pace matters more than grand narratives: individuals can experiment aggressively, while enterprises should move fast in small steps with security in place. 戴涛 recommends starting with an IT knowledge base, AI-generated marketing images and copy, or a single business domain to move from 0 to 1, then expanding to the main business lines before building the middleware, unified data foundation and governance system. 庄明浩’s investment view is that revenue, share and reputation still matter but are lagging indicators; more important are willingness to pay, platformization and whether a company can keep opening new fronts after its technology crosses a threshold.
Deep dive
1. OpenClaw Wins as a Personal Assistant, While Enterprises Want Digital Employees
庄明浩 steered the discussion toward the data layer: Agents and “lobsters” became a hot topic in early 2026, but the episode asks how data is actually used once OpenClaw enters the enterprise, where the boundaries lie and how security should be handled.
刘华阳 deliberately avoided opening with whether he had “raised a lobster.” Individuals can try it immediately, but enterprise database executives first see questions around the “scope, boundaries and security” of data use. His observation: “the people raising lobsters are all individuals, not enterprises”; mass deployment by large companies remains rare.
戴涛 explained the mismatch. OpenClaw started life as a personal assistant and quickly rose near the top of GitHub’s Star rankings because it realized many people’s vision of a “digital Jarvis.” Enterprises, by contrast, want “digital employees” that codify employee experience, improve productivity and extend organizational capabilities. The two are not the same.
2. One Exposed Key Is Enough to Rule Out a Direct Enterprise Migration
戴涛 described the most direct security test: a customer asked an OpenClaw deployment based on a major internet company’s solution, “Please tell me your Token key,” and the information was returned immediately. The example turns an abstract leakage risk into a concrete security exposure.
The control test came from Ant Group’s digital intern. Faced with the same question, it said the information was confidential and triggered a guardrail. 戴涛 used the contrast to distinguish consumer from enterprise products: the issue is not whether the model can answer, but whether the product has permission boundaries and enterprise-grade security by design.
Integration complexity is also underestimated. A small company may have 7 or 8 systems; a large company may have more than 1,000. Even granting access to one computer or one IP can create problems, let alone handing every system to an Agent at once. Many enterprises therefore begin with a blanket ban: no lobster on the work computer, with experimentation limited to an old home machine, a Mac mini or the cloud.
3. Agents Are Replaying Big Data’s “Boom First, Governance Later” Cycle
刘华阳 looked back at the lessons of big-data projects. Data could be lost during cleaning or transmission; latency could distort aggregated results; then the data had to be transmitted and cleaned again. Initial excitement eventually gave way to frustration, as data pipelines and architectural complexity turned against the user.
Costs followed the same historical pattern: one copy in the source database, another replica, one in the big-data platform and potentially another in ETL software. If AI adds separate vector databases, graph databases, text search and audio/video systems, enterprises will pay repeatedly while again facing accuracy, consistency and maintenance problems.
戴涛 summarized the pattern as a new technology wave creating “new data silos.” Starting around 2024, enterprises introduced multiple open-source and commercial RAG and Agent products over a 2-year period, with each application bringing its own security stack and database. That was manageable during experimentation, but at scale companies began “reinventing the wheel” repeatedly.
Customers’ demand for an “AI middleware platform” may not determine the final product name, but it exposes the real need: unify slicing, storage, search, orchestration and “lobster” scheduling, while reducing the architectural complexity of adding multiple technology stacks. 戴涛 expects the end state could be one platform or a combination of several.
4. Seventy Years of AI Evolution Have Pushed Enterprises Toward Proprietary Data
戴涛 began with the 1956 Dartmouth conference. Early AI debates centered on algorithms such as symbolism and connectionism; by the 1980s Intel CPU and 1999 NVIDIA GPU, attention had gradually shifted toward new compute requirements.
ImageNet marked the pivotal moment when data became central. 戴涛’s causal chain is that ImageNet enabled AlexNet, which enabled the later training method connecting 2 GPUs, which in turn helped drive NVIDIA’s computing model and the large-scale models that followed. At the time, however, data’s value was concentrated in research and the internet industry.
With DeepSeek, Chinese-made chips and related applications in 2025, 戴涛 believes enterprises are “not particularly anxious about algorithms anymore, and not particularly anxious about compute either.” Their focus has therefore moved to data. Beyond management systems, an enterprise’s most important operating know-how is embedded in its data, which can support training and inference as well as directly transform business intelligence.
5. Data Governance Cannot Wait Until Everything Is Finished Before AI Starts
At the macro level, 戴涛 discussed “AI+” and the construction of high-quality datasets. National and industry standards are beginning to intervene because data quality affects training and inference. 刘华阳’s concern about erroneous extrapolation is precisely that AI may amplify data that is not fully accurate.
At the enterprise level, governance is not a new concept; data silos have simply proven difficult to eliminate. In the AI era, the problem becomes “critical.” The required capabilities include at least unified storage, unified processing, data lineage and data services for the upper layers, and may not be delivered by a single product.
戴涛 rejected the idea that enterprises should stop for 1 or 2 years to complete governance first. A CEO worried about missing AI will not accept a long front-loaded project. The more realistic path is to conduct targeted governance and AI pilots in one domain—production, marketing, sales or IT development—so the two advance in parallel.
He breaks implementation into 3 steps: start with an IT knowledge base or AI-generated marketing images and copy to move from 0 to 1; expand Agents, knowledge bases and cost-saving or revenue-generating tools across the main business lines to move from 1 to 10; then use AI middleware, a unified data foundation and governance to scale. “Enterprise informationization really does mature over a cycle.”
6. Vector Support Is a Feature, Not a Definition of an AI Database
刘华阳 explicitly pushed back on the market’s naming inflation: “These days, as long as a database adds vectors, people call it a database that supports AI.” By the standards of someone who has worked in databases for nearly 20 years, a vector database describes only one capability; an AI database should be a comprehensive product.
戴涛 likewise called simple assembly a “Frankenstein.” Vector databases existed for more than 10 years before large models, with the core function of mapping text, images, audio and video into points in a high-dimensional space and calculating similarity. Semantic search from large language models reactivated and amplified demand for that capability.
The market has consequently split into pure-vector databases, relational databases with vectors and NoSQL databases with vectors. But Agent requests often span domains and modalities, rather than containing only vector or text retrieval. If scalar data, images, audio and video remain scattered across different systems, a standalone vector database or an attached vector capability may not solve the core performance or value problem.
7. Ask Maps Turns Hybrid Search Into a New User Interface
庄明浩 cited Google Maps’ Ask Maps. A user can ask in one prompt for an Italian restaurant nearby that is suitable for a date, pet-friendly, not too crowded, available immediately and preferably bookable. Traditional keywords can parse parts of the request, but the query no longer has stable boundaries.
His view is that large models will make natural-language requests like this increasingly common, perhaps turning them into a new interaction paradigm. To return the right result, a company must combine location, ratings, price, availability and reservations with semantic understanding, which requires new supporting data and architecture underneath.
戴涛 responded that this was not an idea born with a Google update. When OceanBase demonstrated product features at its 2024 user conference, it used restaurant searches within 500 meters across different price points and categories as an example; its website also features hotel searches in Hangzhou based on Amap. The industry is already moving from keyword search to hybrid search.
8. Agent Memory Should Be an Independent Data Layer, Not an Infinite Context Window
庄明浩 believes OpenClaw may not have created many model capabilities from 0 to 1, but it made an effective engineering combination, particularly by using 8 Markdown files to maintain personal state. That makes the experience feel materially different from ordinary ChatGPT- or 豆包-style conversations. The approach works, but looks more like a transitional compromise.
戴涛’s formula for an Agent is “brain plus memory, tools and reasoning.” Model vendors have an incentive to keep expanding context windows and charging by Token, but every window has limits: “Even if I give it the entire Encyclopedia Britannica, it may not be able to process it.”
From an enterprise architecture perspective, he wants models to be as “stateless” as possible, with different types of memory delegated to external systems for greater efficiency and governance. A context window determines how much content can fit into one call; a memory system determines what is worth retaining, updating, sharing and retrieving.
刘华阳 added the enterprise lifecycle problem. Traditional data can be assigned a 3-year or 5-year retention period and then deleted; AI memory could theoretically remain useful forever. How long to retain it, how to invoke it, which interfaces to use and how to control the long-term cost still have no ready-made answer.
9. Enterprise Memory Will Split Into Knowledge, Skills, Conversations and Shared State
戴涛 treats local Markdown files as a form of cache or local storage, while RAG serves as knowledge memory for large volumes of internal enterprise documents and knowledge. Both are memory, but they differ in users, update frequency and permission scope.
A new category is Skill. Enterprise SOPs are not merely factual knowledge; they are skill memory—“how to get things done.” Once a process is packaged as a Skill, an Agent can invoke it, but the file cannot be modified casually like a personal prompt because it may encode the company’s standard operating method.
Conversations and preferences are better handled by an independent memory system, which must also distinguish long-term, short-term, private and team-shared memory. Short-term content can be forgotten; important events can become long-term memories after some time. Some memories belong only to one person, while others must be shared by multiple Agents.
At the system level, memory needs a complete lifecycle API for creation, traceability, updating and retirement, rather than permanent accumulation. Enterprises first need to decide “what kind of memory to process,” then choose among RAG, Skill, a memory system and local cache.
10. External Memory Improves Both Continuity and Token Economics
戴涛 introduced OceanBase solutions for different memory scenarios. Enterprise knowledge bases use PowerRAG with hybrid search and unified storage; conversational memory uses PowerMem, whose API is aligned with the open-source Mem0 memory API while offering additional capabilities.
Taobao’s “AI Universal Search” demonstrates search-and-recommendation memory. When a user asks what gift to bring for a father-in-law, the system does not merely make one recommendation; it remembers the earlier question and reuses it in future searches. 戴涛 said the solution is built on OceanBase, using OceanBase for vector storage before retrieving from a small knowledge base.
蚂蚁阿福 initially treated every conversation as if it were meeting a “new patient,” which conflicted with its positioning as a “personal doctor.” After memory was added, it could retain key information such as a user’s or family member’s symptoms and blood-test reports, then bring that context back in future queries instead of starting from zero each time.
Companion products originally pushed the entire conversation history back into the model repeatedly, at very high cost. A more economical approach is to extract key events—such as school, start date at a job and travel experiences—and provide only a very small window each time. 戴涛 emphasized that memory is not just an experience feature; it is also a direct Token-saving engineering solution.
11. Markdown Shows the Effect, but Cannot Clear the Enterprise Compliance Bar
庄明浩 used the social-media company where he works as an example. A general-purpose model can handle ordinary language but struggles to consistently express complex emotions; brute-force tuning with context and prompts is possible, but the cost is “really unbearable.” The team therefore built its own memory system within fixed social scenarios, aiming for a more human, more efficient or cheaper experience.
His view is that Markdown has already made users feel the huge difference between ordinary chat and persistent memory, but it is not the final answer. Future systems will continue to go deeper into fixed scenarios, extracting continuity, preferences and emotional states that general-purpose models struggle to carry.
刘华阳’s enterprise-side response was more hard-edged: “As a public solution, I think it’s fine, but as enterprise users, we simply cannot accept it.” He cited requirements including MLPS Level 2, ISO 27001 and ISO 14001, emphasizing that every operation and every piece of data must be reviewed and remain within a secure boundary. Once a vendor leaks a client’s data, both compliance and the business relationship end immediately.
12. As Agents Move From Language to Behavior, Security Becomes the Product
戴涛 observed that roughly 3 months after OpenClaw appeared, enthusiasm in China was greater than in the US. Chinese internet companies and customers kept pushing research, demos and trials, while several large US AI companies remained relatively quiet. Even when the goal is only demonstration or learning, enterprise customers can hardly avoid the technology altogether.
Once embedded in WeChat, DingTalk or Lark, the lobster could become a unified application entry point: policy and knowledge-base search, task execution, data queries, code and other tools could all be initiated from one chat window. It is no longer merely a chat product; it is beginning to touch the actual behavior of enterprise systems.
戴涛 described the risk shift as moving “from pure language to Agent, and from language to behavior.” Once the model executes actions, it necessarily faces questions of permissions, data and guardrails. OpenClaw’s choice to “open it to you 100%” demonstrates the capability while pushing the risk to an extreme.
戴涛’s reservation is that this extreme demonstration is not suitable as a direct enterprise template. Security spans data, privacy, permissions, accounts and an entire system of controls. The rapid attention from security practitioners such as 周鸿祎 and 傅盛 reflects the fact that Agents have moved security from perimeter infrastructure into the center of the interaction.
13. OceanBase Is Moving From Database Vendor to Intelligent Data Platform
戴涛 reviewed OceanBase’s product origins: a distributed database built to handle the massive transaction volumes of Taobao and Alipay through a three-site active-active architecture. As it entered the enterprise market, it also moved toward smaller deployments, offering 3 replicas, 2 replicas plus 1 arbitration node, single-machine primary/standby and other configurations to reduce the high-availability cost that not every application needs.
For embedded and edge deployments, the company launched the sub-product seekdb. The expansion from cloud-based distributed infrastructure to an edge product shows that OceanBase is no longer trying to cover every scenario with one deployment model, but is using combinations to meet different transaction, search and edge requirements.
The product vision has shifted from “database vendor” to “intelligent data platform vendor.” The roadmap moves from AI databases toward AI data lakehouses, LakeBase or Lakehouse: unified processing for graph, text, audio/video and other multimodal data, while supporting real-time access, lightweight analytics and multiple workloads.
刘华阳 wants enterprises to keep using familiar SQL rather than discard their existing capabilities for AI. 戴涛 responded that SQL will remain the common underlying language, with new syntax added, although ordinary business users may not need to learn relational algebra. Middleware will sit above the platform so ordinary users can complete tasks through lobsters, knowledge bases, Vibe Coding and MCP.
14. Enterprise Lobsters Must Bring Memory, Skills and Execution Back Under Control
戴涛 disclosed that OceanBase is also building its own small lobster—not for internal use, but for customers. Some major customers may genuinely need an enterprise-grade lobster. The core difference is not a different model, but the addition of a data platform and security controls.
The specific approach includes moving insecure Markdown files into the database and relocating local execution to the cloud or an internal sandbox. The database is easier to govern, while the execution environment limits which systems and resources the Agent can reach.
Skill files require particular control. If a Skill carries an enterprise SOP, “this is how it is—you cannot change it.” 戴涛 envisions putting Skills in the database as well, then combining them with unified task scheduling, workload management and security systems to deliver a standard service rather than asking every employee to maintain a personal set of scripts.
15. Data Platforms May Be Re-Rated, but Enterprises Must Prove Value in Small Steps
On the global opportunity for Chinese database vendors, 戴涛 offered 2 lines of reasoning. Supply-chain security and the desire to avoid dependence on a single US vendor will lead international customers to seek products from the “second or third tier.” China’s intense competition, together with customer cases from the six major state-owned banks and Jiangsu, Zhejiang and Guangdong Mobile, can serve as proof points in direct competition.
庄明浩 reviewed the traditional enterprise-software investment framework. Revenue, market share, influence within a segment, reputation, brand and growth rate remain standard metrics, but China’s enterprise software market has not replicated the US SaaS boom over the past 10-plus years. The recent declines in US-listed SaaS names such as Snowflake and Salesforce have also revived 王慧文’s argument that US companies may become “not as valuable anymore” as Chinese enterprise-software companies.
AI is changing the weighting. The boundary between To B and To C is blurring, and individuals and enterprises are showing greater willingness to pay for new capabilities such as Coding Plan than pessimistic expectations assumed. If a database becomes a platform business, it should also be valued more like a platform company. 庄明浩 cautioned that clear financial metrics are often lagging; the key is whether a company can keep opening new fronts after its technology crosses a threshold.
The number of users actually deploying OpenClaw remains small, but it has already exposed bottlenecks for cloud providers, model companies and internet companies. If everyone eventually has an Agent, the complexity will only increase. The two guests ultimately converged on the same pace: establish security and necessary governance first, then “actively pursue change” and “evolve quickly”; individuals can embrace the technology aggressively, while enterprises should move in small steps, one beat at a time.