"Agents manage over $39 million in positions" sounds like meaningful scale — so why does the same report also say agents lose to humans by 5x at trading? Aren't these two numbers contradictory?
They're not, because the two numbers describe different categories of task. The $39 million in agent-managed positions is overwhelmingly concentrated in yield optimization — a "narrow objective, stable rules" scenario where moving capital between protocols to chase rates is something agents genuinely do well, which is exactly why that figure could grow to meaningful scale. The 5x underperformance comes from a separate benchmark — open-ended trading contests measuring the ability to judge entry and exit timing and bear directional risk. In other words, the scale of capital agents currently manage has grown precisely because they're deployed on the task type they're good at — that's not contradictory evidence about overall agent capability, it actually confirms the core logic that task type determines performance.
Why is open-ended trading specifically difficult for agents? Is the bottleneck insufficient information, or insufficient compute speed?
It's neither information volume nor compute speed — today's agents access market data and compute far faster than any human. The real difficulty is that open-ended trading requires making judgment calls where there's no single correct answer, and that kind of judgment is often built on reading the current market narrative, crowd psychology, or an intuitive comparison of "how is this situation similar to or different from a past one" — things that are extremely difficult to reduce cleanly into rules or parameters. Yield optimization is simple precisely because its objective can be reduced to maximizing a single number (annualized return); the "objective" in open-ended trading itself contains a large component of judgment that can't be fully quantified, and that's precisely the part current agent architectures struggle most to replicate.
If I want to personally verify a yield figure a DeFAI product claims (like the 9.75% annualized figure mentioned in the report), is there an actual way to check it?
Things you can check include: what time window the figure was measured over (short-term high yields often don't represent a long-term average), whether the calculation nets out gas costs, Slippage, and any platform fees before arriving at a net yield, versus reporting a gross figure only, and whether that agent's historical trading record is publicly verifiable on-chain rather than relying solely on numbers a platform's own dashboard reports. Another often-overlooked check: whether the agent is still actually operating. If an agent once hit an impressive number but has since stopped operating or scaled down substantially, that historical figure's relevance needs to be re-evaluated — it shouldn't be cited as something currently happening.
As an ordinary DeFAI user, how does knowing "agents are strong at yield optimization but weak at open-ended trading" actually apply to my own decisions?
In practice, you can use that distinction as a screening filter: if a DeFAI product's advertised function is moving capital between a fixed set of protocols according to clear rules to chase yield, that category currently has relatively solid data supporting its viability, and the main risk comes from the underlying protocols rather than the agent's judgment. But if a product's advertised function is an open-ended trading feature — the agent autonomously judging market direction and timing entries and exits — current data shows that's exactly where agents perform weakest and where risk is hardest to estimate, which calls for a much higher degree of scrutiny toward any historical return claims tied to it. Put simply: an agent's degree of "autonomy" and its degree of "trustworthiness" are not the same thing at this stage.
A report published by crypto venture firm DWF Ventures on April 16, 2026 offers the most concrete quantitative picture yet of how DeFAI agents actually perform: autonomous agents now drive more than 19% of on-chain activity, with over $39 million in Total Value Locked across agent-managed positions. But the same report also surfaces a contrasting figure that's easy to lose under marketing narratives — in open-ended trading tasks, the best-performing human traders beat the best-performing agents by more than five times.
The report quotes Xin Yi Lim, a senior investment analyst at DWF Labs: "Agents thrive when the objective is narrow and the parameters don't move often." That line draws a clear dividing line. Yield optimization — moving stablecoins between lending platforms to chase the currently highest rate — is a textbook example of a "narrow objective, stable rules" scenario: the comparable options are limited, the success metric is unambiguous (annualized yield), and there's no ambiguous situation to interpret. The report notes that Giza's ARMA agent, while it was operating, achieved a 9.75% annualized return, outperforming the yields other DeFi protocols like Aave and Morpho were offering at the time.
Open-ended trading — judging entry timing, bearing directional risk, adjusting positions as markets swing — requires a fundamentally different Skill: making judgment calls with no single correct answer, under conditions that keep shifting and are ambiguously defined. The report cites a stock trading contest where the top-performing human trader outperformed the top-performing agent by more than 5x. In a separate AI model trading contest (nof1), only 3 of 7 participating models ended up profitable per trade. Coinbase CEO Brian Armstrong and 0G Labs chief growth officer Aytunc Yildizli are both quoted in the report, converging on the same observation: agents currently execute reasonably well, but their judgment still lags far behind human judgment.
The root of this gap isn't that agents compute too slowly — it's that open-ended trading inherently involves rules that themselves keep changing. The same strategy might require a completely different response depending on market structure and liquidity conditions at the time. In a yield optimization task, the "rules" are relatively stable: which protocol offers the higher rate, what the gas cost is — these are clearly quantifiable variables that can be tracked continuously. In open-ended trading, the "rules" themselves are the uncertainty — nobody can tell an agent in advance which logic the market will follow next. The report estimates it will take another 5 to 7 years before agent trading volume genuinely rivals human volume.
This data gives anyone evaluating a DeFAI product a more useful screening standard than a blanket "AI is great" or "AI isn't ready" verdict: the question to ask about a given DeFAI agent isn't whether it's strong "overall" — it's whether the task it was designed to handle falls into the "narrow objective, stable rules" category or the "open-ended judgment" category. The former currently has real data backing solid agent performance; the latter remains a domain where humans still hold the advantage. Judging the entire DeFAI category by a task type agents aren't good at, or over-extrapolating from a task type they are good at to assume equal reliability elsewhere, are both common misjudgments worth avoiding.