
From Qianwen Work to WorkBuddy: Harness as the Execution Middleware Becomes a "Must-Have" in the Battle Among Tech Giants
Leading vendors are intensively integrating the Agent middleware layer, Harness, making it a critical component in the competition among tech giants. Driven by the need for long-horizon execution and deterministic model iteration, Harness lowers barriers to entry and accumulates high-quality data, fueling a flywheel of TAM expansion and product enhancement. Alibaba, Tencent, OpenAI, and others are strengthening this layer; we recommend paying close attention to vendors with mature Harness capabilities
Recently, leading vendors have been intensively integrating the Agent intermediate product layer, with Harness transitioning from an optional component to a necessity. The reasons are twofold. First, there is rigid demand from the customer side: as Agents move toward long-horizon execution and Multi-Agent collaboration, costs related to state maintenance, error propagation, and Tokens rise rapidly. Strong single-step model capability does not equate to stable delivery, making Harness a prerequisite for end-to-end tasks. Second, there is a demand for determinism on the model side: amid converging capabilities, accelerated iteration, and intensified competition, closed-loop business models and high-quality data have become scarce sources of certainty for model iteration. Harness provides both: on one hand, it uses natural language interfaces to push Agents from narrow coding scenarios to broader white-collar applications, expanding the addressable market from approximately $200 billion to $1.5 trillion; on the other hand, it accumulates scarce corpora such as long-horizon task trajectories and correction feedback, potentially forming a "TAM expansion—data feedback—product enhancement" flywheel that feeds back into model iteration. Harness currently bridges and binds model supply with customer demand: it connects and orchestrates multiple models upstream, while supporting development, research, and general office scenarios downstream. It serves as the middleware layer that transforms underlying intelligence into stable delivery. We recommend prioritizing vendors that already possess mature Harness capabilities.
▍ Report Background: Leading Vendors Intensify Integration of Agent Products, Shifting Competition from Single-Point Capabilities to Unified Entrances and Task Delivery.
Recently, Alibaba integrated QoderWork, Wukong, and MuleRun into Qianwen Office; Tencent promoted the convergence of the WorkBuddy and QClaw teams; OpenAI further unified the entrances for ChatGPT and Codex; and xAI strengthened the combination of models and development workflows through Cursor. While these moves vary in form, they essentially all reinforce the Harness layer situated between models and user scenarios. We believe that Agents are evolving from "answering questions" to "delivering results." While models determine the upper limit of intelligence, Harness determines whether model capabilities can be stably, controllably, and cost-effectively transformed into tangible outcomes. Its core value lies, on one hand, in lowering usage thresholds and expanding the general white-collar TAM; on the other hand, in accumulating high-quality long-horizon task trajectories and feedback data to feed back into continuous model and product optimization.
▍ What is Harness: The Agent Intermediate Product Layer Connecting Model Supply and Customer Demand.
From a technical evolution perspective, Agent engineering has moved from Prompt optimization for expression and Context construction for information environments, further toward Harness-driven process control and stable execution, converting production trajectories into continuous optimization signals via Loops. Models handle probabilistic reasoning, while Harness assumes responsibilities for Context, Memory, tool invocation, permission control, state management, and result verification. From a product perspective, the boundaries of Harness extend further to a broad intermediate layer encompassing multi-model routing, Multi-Agent orchestration, Skills and Connectors, Runtime and Sandbox, task entrances, and result delivery. We believe that Harness connects and organizes multiple models upstream, while supporting research, development, data analysis, and general office needs downstream, serving as a key ecological layer that encapsulates underlying intelligence into end-to-end task products.
▍ Why Harness is Needed: As Agents Move from Usable Capabilities to Scalable Delivery, Execution Systems Become Key to Competition.
The direction of global model iteration has clearly converged toward Computer Use, Tool Use, Agentic Coding, and long-horizon tasks. User evaluation criteria have also upgraded from "how good the answer is" to "whether the task can be completed and results delivered." However, single-step model capability does not equal end-to-end delivery capability: according to METR, the success rate of Claude Mythos Preview drops to 50% when executing tasks lasting approximately 17 hours; according to Anthropic, Token consumption for Multi-Agent systems is about 15 times that of ordinary Chat, with bottlenecks in state maintenance, error propagation, and synchronization rising as scale increases. According to Microsoft, CodeAct Harness reduced execution time by 52.4% and Token consumption by 63.9% by optimizing tool orchestration, further illustrating that Harness determines whether multiple Agents can form an efficient organization. Meanwhile, models have crossed the baseline usability and cost thresholds for knowledge work. According to Analysys, monthly visits to desktop office Agents in China increased from over 20 million in March 2026 to over 60 million in June. The industry contradiction is shifting from "whether it can be used" to "how to achieve stable, low-cost, scalable delivery."
▍ Recommendation to Prioritize Harness: From Model Distribution to User Entrances, TAM Expansion and Data Feedback Strengthen the Positive Flywheel.
As open-source models accelerate their catch-up and the substitutability of underlying capabilities increases, product competition is likely to shift from "binding to the single strongest model" to "continuously invoking the model best suited for the current task," thereby enhancing the value of Harness in model routing, user entrances, and task distribution. Products like WorkBuddy opening access to multiple models reflect that leading vendors are trading off model exclusivity for larger user and scenario TAMs. User relationships, task trajectories, tool invocations, and correction feedback accumulate at the Harness layer, forming a flywheel of "open models—expand TAM—secure entrance—data feedback—product enhancement," where high-quality data continuously feeds back into model capability improvements. Meanwhile, according to our calculations, replacing professional IDEs with natural language office entrances is expected to expand Agents from the ~$200 billion AI Coding market to broader knowledge work, bringing the addressable space for AI suppliers to approximately $1.5 trillion.
▍ Investment Strategy:
The focal point of Agent competition is shifting from model capabilities to task delivery. The convergence of capabilities at the model layer and the shortening lead window are moving industry certainty toward the intermediate layer. Vendors with mature Harness capabilities simultaneously hold two scarce assets: first, a closed-loop business model spanning unified entrances, task execution, and result delivery; second, high-quality data such as long-horizon task trajectories and correction feedback that is difficult to obtain externally. These two assets reinforce each other, making the flywheel of first-movers harder to catch up with. Therefore, we recommend prioritizing platforms that have successfully implemented Harness product forms and possess synergistic capabilities in models, cloud services, and high-frequency office entrances.
Risk Warning and Disclaimer
The market carries risks; investment requires caution. This article does not constitute personal investment advice, nor does it take into account the specific investment objectives, financial status, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article align with their specific circumstances. Investors bear full responsibility for their own decisions.
