Are autonomous agents already able to do real work unsupervised?
34 recorded positions from 31 people, first said Sep 27, 2023. They do not agree — the readings below are what each one actually argued.
Narrow scoped jobs succeed open ended scope fails
Aatish Nayak · Apr 11, 2025
For complex knowledge work, agents will need extensive human context and will have to ask users many questions rather than operating as fully automated black boxes
People don't trust black boxes, and complex work depends on uniquely human elements like context and emotion
Scope: simple tasks like booking a flight or ordering dinner can be fully delegated; applies to architecture, legal, tax and enterprise knowledge work
53:14 20Product: How Scale AI and Harvey Build Product | Why PMs Are Wrong: They are not the CEOs of the Product | How to do Pre and Post Mortems Effectively and How to Nail PRDs | The Future of Product Management in a World of AI with Aatish Nayak
Varun Mohan · Jun 2, 2025
In the short term, async remote agents will only handle tasks easy enough to be completed correctly, while complex work stays in a local agent like Windsurf
Asynchronous output that takes hours faces very high quality expectations (near 99% correct), and correcting a 90%-correct pull request still requires understanding the wrong 10%, so the feedback loop is too slow for complex, iteration-heavy tasks
Scope: short term
37:08 20VC: Windsurf Founder on Will Model Companies Own the App Layer | Why Moats Do Not Exist in a World of AI | Why the Notion of Single Person $BN Companies is BS | Lovable vs Bolt & Cursor vs Windsurf: How Does it All End with Varun Mohan
Anish Acharya · Feb 9, 2026 · hedged
The agent-maximalist view — fully autonomous agents doing everything over long time horizons — is ahead of where we actually are; humans must stay in a tight loop with models
Exception handling requires humans, and models are only as good as instructions, which in practice are frustratingly vague, just as they are when managing a team
Scope: may get there someday
55:59 20VC: Is SaaS Dead in a World of AI | Do Margins Matter Anymore | Is Triple, Triple, Double, Double Dead Today? | Who Wins the Dev Market: Cursor or Claude Code | Why We Are Not in an AI Bubble with Anish Acharya @ a16z
Avishai Abrahami · Jul 13, 2026
The gap is between open-ended AI and narrowly scoped AI: pointed at a complicated company's full customer support it fails, but built with the right skills and tools for a specific job like managing orders or schedules it can be incredibly efficient
Capabilities like voice conversation that were impossible four or five years ago are now real, showing what focused AI can deliver
35:11 20VC: Wix's Founder on What Wall St Gets Wrong About AI and Wix | Will Base44 Win the Vibe Coding Wars | The Truth About the Economics of Vibe-Coding | The Buyback Disaster: Lessons Learned with Avishai Abrahami
Coding is the only domain at parity so far
Anjney Midha · Apr 14, 2026
Recursive self-improvement is genuinely happening in coding-adjacent tasks like automating data analysis and cleaning, but it does not extend to bootstrapping physical infrastructure such as a real-world R&D lab
Setting up permitting, raising money and building physical infrastructure is an entirely different kind of execution problem than code automation
Scope: limited to tasks with coding involved
9:24 20VC: Anj Midha on Investing $300M into Anthropic | The Early Days of Anthropic & How 21 of 22 VCs Turned it Down | The Four Bottlenecks to Compute | What the China Has Smashed and Why We Should Be Worried
Aaron Levie · Apr 20, 2026
People wrongly assume the outcomes seen in AI coding will quickly transfer to other areas of knowledge work
Coding has idiosyncrasies, and the rest of work happens differently in ways that don't map over
46:10 20VC: Everyone is Wrong; We Will Have More Developers in Five Years | Why Frontier Labs Will Be Way More Valuable Than They Are Today | Are SaaS Companies Cooked: Which Thrive & Which Die with Aaron Levie, Founder at Box
Julien Bek · Aug 24, 2026 · hedged
Agents are currently only good at selecting tools in coding, and agent-driven tool selection in other functions is much further out on the adoption curve
Coding is the one area where agentic applications have reached human parity; other functions haven't
Scope: framed as a position on the adoption curve; 'probably'
53:14 20VC: Inside Sequoia's Investment Committee: Lessons from Don Valentine, Doug Leone and Alfred Lin | How the SpaceX and Citadel Deals Went Down | What Sequoia Specifically Looks for in Founders with Julien Bek
Humans must remain validators ai does retrieval
Christian Lanng · Sep 27, 2023
The right design is to keep humans in the loop as validators while AI does retrieval and correlation, putting the human brain where it is best rather than turning humans into robots
AI can gather and correlate information; humans should spend their judgment on validation instead of navigating bad UIs
53:48 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng
Joelle Pineau · Nov 3, 2025
Human guidance of machine behavior is an enduring partnership, not a temporary training phase
What changes over time is the division of labor — which information the AI supplies versus what humans must supply as a complement — not the existence of the human guidance role
34:09 20VC: Cohere's Chief AI Officer on Why Scaling Laws Will Continue | Whether You Can Buy Success in AI with Talent Acquisitions | The Future of Synthetic Data & What It Means for Models | Why AI Coding is Akin to Image Generation in 2015 with Joelle Pineau
Cliff Weitzman · May 9, 2026
In a world where software engineering and design are commoditized, QA remains a human job — AI coding tools cannot QA themselves to perfection
Coding agents fail across different devices, phones, and degraded network conditions; even good engineers and designers rarely put themselves in the user's shoes and test edge cases
Scope: especially across devices and edge cases like Wi-Fi dropping
58:35 20VC: What I Learned from 100 of the Best CEOs in the World | What I Learned from Staying with Mr Beast for 3 Weeks | How We Will Spend More on Tokens than Salaries with Cliff Weitzman, Speechify
Capability is underestimated retest constantly
Raaz Herzberg · Dec 12, 2025
AI capability is improving so fast that leaders must be willing to change their minds daily about what is possible
Things impossible six months ago are now doable — brand/design generation on your own guidelines went from bad to excellent within a year, and new Gemini is far better at specific security errors than any prior model
62:02 20Growth: How Wiz Built a $30BN Brand in Enterprise | What Worked vs What Was a Mega Failure: Lessons Learned | Why Marketers Make the Worst CMOs & What To Look for in Growth with Raaz Herzberg
Winston Weinberg · Jan 19, 2026
GPT-3-era models were already good enough at legal work that practicing attorneys would send the answers straight to clients
They ran Reddit r/legaladvice questions through a chain-of-thought product and blind-tested the outputs on landlord attorneys, who said 86 out of 100 answers were ready to send
Scope: based on consumer-level landlord/tenant legal questions; blind test where attorneys weren't told AI was involved
60:27 20VC: How Model Performance is Plateauing | Two Key Rules for Effective Deal-Making | Company Building Lessons from Keith Rabois, Brian Halligan and Pat Grady | Why Enterprise AI Adoption is Years Off with Harvey CEO Winston Weinberg
Fred Turner · Jul 18, 2026
Most people severely underestimate current model capabilities because they last tried them a while ago; you should relentlessly retest them and plan for where they're going
Models are moving so fast they're much better than even six months ago
64:01 20VC: $5BN in Revenue, 7 to 7,000 Employees in 9 Months, 206,000 Tests in a Single Day: The Craziest Story in Startups: Curative with Fred Turner
Agents already write production code market hasnt registered it
Mike Krieger · Mar 3, 2025
Agentic coding tools are especially valuable for end-to-end product engineering workflows and for navigating unfamiliar codebases
Product engineering spans backend, frontend and translation submissions in one loop; Mike had never opened Anthropic's codebase yet Claude Code found the right files and made edits for two pull requests
Scope: acknowledges not everybody is in his situation
42:08 20VC: Anthropic CPO Mike Krieger: Where Will Value Be Created in a World of AI | Have Foundation Models Commoditized | When Do Model Providers Become Application Providers | What Anthropic Learned from Deepseek
Jerry Murdock · Feb 28, 2026
Autonomous agents actually writing production code at startups is the most significant current development and is not yet visible to the market
It's only about two months old and the companies doing it have been at it for two to six weeks, so the marketplace hasn't registered it
Scope: based on his own portfolio companies
6:40 20VC: Why Cursor is Dead | An AI Tsunami is Coming & You Need to Prepare | Systems of Record Become Valueless Databases with Agents | Is This The End of Tech Private Equity with Jerry Murdock, Co-Founder of Insight Partners
Tobias Lütke · May 4, 2026
December was an inflection point for AI-assisted engineering — the arrival of Opus changed everything
Many of Shopify's best engineers have not written code themselves since December, and the AI-generated share of code is well over 50% and rising fast
54:57 20VC: Shopify's Tobi Lütke on How AI is a Scapegoat for Mass Layoffs & What Will Labour Markets Be in the Future | Why We Need More Scrutiny on Charitable Giving, Governments are Bad at What They Do and Trump Derangement Syndrome in Canada
Also on the record
Fred Turner · Jul 18, 2026
Empowering agents to act on your behalf — signing contracts, booking travel with your card — is something that seems insane today but will be obvious in ten years, though trust will take time to build
Internally it took many rounds of convincing to let an agent click sign in DocuSign on a legally binding contract, so consumer trust will lag similarly
66:05 Capability arrives before trust in delegated authority
Kieran Flanagan · Jul 11, 2025
Autonomous multi-task agents are the biggest hype in AI go-to-market tooling; today's AI is not reliable enough to guarantee an agent completes multi-step goals consistently
The agents he has used internally and seen at other companies are inconsistent and unreliable, and people are often unclear that an agent requires data, tools and context
15:34 Multi task autonomous agents are overhyped and unreliable today
Nikesh Arora · Jun 22, 2026
Capturing AI's real benefit requires letting AI do roughly 80% of the judgment work in a workflow, which means relinquishing human control
Existing SaaS workflows are hard-coded containers with defined inputs and outputs and no intelligence; only if AI makes the judgments (e.g. selecting candidates and generating interview questions) does the process become fundamentally more intelligent
12:27 Benefit requires ceding judgment not just execution
Matt Clifford · Jul 1, 2024
Reliable agency would be the needle-moving capability shift in the next model generation, because current GPT-4 agents produce great demos but are not robust or reliable
People building agents on GPT-4 report you can cherry-pick impressive runs out of a hundred, but robustness is a huge problem; models good enough for much more reliable agents would be a qualitative change
24:26 Reliable agency not demo cherry picking is the next needed capability leap
Jerry Murdock · Feb 28, 2026
The shocking development is that autonomous agents actually work unsupervised — the agent has gone from assistant to employee
OpenClaw runs on its own without needing review, which is why developers are so passionate about it
30:54 Unsupervised operation turns the agent into an employee
Adam Foroughi · Apr 27, 2026 · hedged
Because agents can already recursively improve the code and products available to them, the rate at which imagination becomes product will keep accelerating and the future will be far more productive than the present
Product rollouts are appearing daily and people who know how to deploy armies of agents, most obviously for coding, are compounding that output
60:48 Recursive agent self improvement compounds output
Martin Casado · Jul 28, 2025
AI is unlike prior general-purpose technologies because it currently requires a human handler, so displacement works differently than with electricity
Models are so unpredictable that all monetized use cases have a human on the other side — coding has a professional coder, creative work has a creator
46:03 Current ai uniquely requires a human handler unlike prior general purpose tech
Daniel Dines · Dec 18, 2024
Agents are idiot savants — sometimes extremely smart, sometimes extremely dumb — and there is currently no way to tell which case you are in
18:01 Agents are idiot savants with no reliable way to predict which mode youre in
Victor Riparbelli · Jan 15, 2025
We are not near being able to automate real software building end-to-end — you cannot yet ask a model to build the next social network and get it — because production software is complex and involves humans, customers and feedback loops, not just code.
Most software in production is highly complex and social; the last two years have produced cool demos (a Tetris game in the browser) but not mass-scale production automation.
22:28 Production software building is too complex and social for current end to end automation
Nick Frosst · Sep 1, 2025 · speculative
By 2026 you will be able to tell an application 'file my expenses' and the model will reliably find the policy and the receipts and complete the whole task, and this becomes a ubiquitous way of using a computer
The capability sounds close, but making it actually work reliably is the hard part — most people and companies have no such experience today, so ubiquity would be a radical change
60:05 Expense filing style tasks become reliably automatable by 2026
Andrew Ng · Nov 17, 2025
Useful AI agents are not a decade away — useful agentic workflows already exist today
AI Fund has built agentic workflows for tasks that were previously impossible, e.g. tariff compliance (Gaia Dynamics), a medical assistant in India, and legal document processing (Callitus); hyperscalers and large businesses also run internal workflows they could not do without AI agents
34:39 Useful agentic workflows are deployed today not years away
Max Junestrand · Jan 26, 2026
We are heading toward a world where people kick off long-running agent tasks before bed and wake up to completed work, but we are not there yet
No task in Legora currently takes twelve hours to run, but once models can be put in loops that improve with each iteration, long-running tasks become viable; deep research already gave a taste of this
18:27 Long running overnight agent tasks are coming but not yet
Sarah Tavel · May 6, 2024
Some employee work products can already be fully automated by AI today, though automation readiness is a spectrum rather than all-or-nothing
She has met companies automating HR ops, recruiting and sales work; thinking of it as 'unbundling the employee' into discrete work products, some of which are automatable now
14:05 Unbundling employee work into discrete products some already fully automatable now
Harry Stebbings · Apr 7, 2026
When configuration files change so they no longer configure the agent, an agent set up to perform a task completely falls over
11:29 Agents are brittle to environment changes and collapse when config shifts
Ryan Petersen · Jun 20, 2026
Agents can automate enterprise logistics business logic end-to-end in a way that RPA never could, replacing unwieldy human-maintained rules engines.
Every enterprise customer has bespoke rules and data formats, producing if-then sprawl that humans end up managing; RPA went part of the way, agents can go all the way.
15:42 Agents absorb bespoke business logic rpa could never reach
Richard Socher · Apr 18, 2025
Action agents that take irreversible actions on the web are in a valley of disillusionment because they don't yet know enough about the individual user
Preferences shift with life circumstances — a student wants the cheap one-stop flight, an employed person with money and less time wants the direct flight — and the AI needs those subtle personal details to be good
16:28 Action agents stall without deep personal preference data
Harry Stebbings · Aug 15, 2026
An AI call grader that stack-ranks companies across variables like founder-market fit and product-market fit is reliable enough that a fund can fully delegate investment prioritization to it
In 12 weeks of running it across over a thousand companies the ranking has never been wrong
71:55 Reliable enough to fully delegate ranking decisions
Matt Swulinski · Aug 15, 2026
You can trust an AI system's decisions when you have done the upfront work of specifying the decision tree, because the output is then based on your own framing
The system's judgments derive from the framing you supplied, so trusting it is trusting your own thinking
72:27 Trust follows from specifying the decision tree yourself
Your assistant can query this graph directly — 34 positions here, 19,646 across the corpus. Add 996.fm over MCP.