Methodology

The Two Hills of AI: Automation vs. Invention

Most enterprise AI strategy is framed around productivity: automate an existing manual process, count the hours saved. Real value — but productivity gains from AI automation are table stakes. If you can automate a process with AI, so can your competitor, usually by buying the same tools. Efficiency everyone can purchase is not a differentiator; it's the new baseline.

The more interesting question: what can AI let us do that was previously impossible?

Invention subsumes automation. The reverse is not true. The invention path, done right, delivers the productivity gains anyway — as a byproduct.

The Two Hills of AI framework was introduced by Sudhir Patavardhan in the original LinkedIn article of the same name, and this analysis practice was first run under the name Two Hills Lab; every report on theAIFolks applies it.

The 2×2

Scope of impact × nature of use

IndividualInstitutional
Invent Individual InventionBrilliant one-off tools built by your best people. Genuine innovation, but it dies when the builder moves on. Institutional CapabilityNew abilities embedded in how the organization itself works. Compounding. Hard to copy. Where durable value lives. Finish here.
Automate Personal ProductivityCopilots, prompt-assisted work. Valuable, but it walks out the door with each employee. Start here. Process Automation at ScaleAutomated organizational workflows. Real savings, replicable by anyone with a budget. Efficiency everyone can buy is the new baseline.

Every AI journey starts in the bottom half — that's where the quick wins are. Durable value lives in one quadrant only: institutional capability. Personal productivity is a fine place to start. Institutional value is the only place worth finishing.

The diagnostic question, for every AI initiative on a roadmap: is this automating something we already do — or inventing something we couldn't do before? Both are legitimate. Only one compounds.

The reference case

Enterprise quality engineering

The framework was built from one lived case: reimagining end-to-end UI test automation at scale.

  • Obvious path (automation): AI generates test scripts from natural language. A legitimate productivity win — but real-world usage showed generated scripts couldn't keep pace with application change. Same constraints, faster output.
  • Different question: where should the intelligence live? Answer: in understanding the intent of a use case and figuring out execution the way a human tester would — not in generating scripts.
  • Result (invention): a platform that reasons about user intent rather than producing fixed scripts. Building for capability rather than a specific tool meant it covered multiple platforms and device types, scaled to a very large number of use cases, and kept its own tests current as the product evolved. Productivity gains followed as a byproduct.

This case sits on the Unsolved Problems index as the calibration entry every domain is measured against.

Applying the framework to a new domain

Five steps, same order, every domain

  1. Map current investment. Where is the money and attention actually going? Classify each major initiative or vendor category into a quadrant. Expected finding in most domains: heavy concentration in the bottom half. That concentration is the story — evidence of a commoditizing baseline.
  2. Identify the binding constraints. What has always made the domain's hard problems hard? Cost of expert attention that doesn't scale, unreadable unstructured data, combinatorial search spaces, feedback loops too slow, coordination costs across silos.
  3. Test the constraint against AI. For each "unsolvable" problem: does AI remove the actual binding constraint, or just accelerate work within it? Producing output faster while the fundamental constraint remains is automation dressed as invention — the script-generation trap.
  4. Find the institutional-capability formulation. Ask "where should the intelligence live?" — usually upstream of where the automation tools sit. What new ability could be embedded in how organizations in this domain work, such that it compounds and is hard to copy?
  5. Falsification check. What evidence would prove this domain is an exception — that automation is the moat here? Actively look for it. Record the outcome either way.

Anti-patterns we watch for — in others' initiatives and in our own analysis: faster output, same constraint (the script-generation trap); tool worship (building for a specific tool instead of a capability); purchased differentiation (claiming moats from vendor products competitors can also buy); and individual invention mistaken for institutional capability (genius one-offs that die with the builder).

Evidence discipline

The P/R/N/V tier system

Every claim used anywhere in a report must first appear in that domain's evidence log — numbered, dated, and tiered. The tier is a trust ranking, not a category, which is why it renders as a single hue stepped by strength rather than four unrelated colors. Each log ends with a "Known gaps" block: the claims we looked for and could not establish. Showing what a report doesn't know is part of the method, not a disclaimer.

P
Primary

Filings, earnings calls, official company statements, court filings, government data. The strongest tier — a party speaking on the record about itself, with consequences for lying. Note: self-interested primary sources (a consortium reporting its own scale) are still flagged in the row notes.

R
Industry report

McKinsey, Bain, BCG, Coresight, Gartner, academic and peer-reviewed work. Methodical, but often gated, sometimes secondary-sourced — the log notes when a figure could not be verified against the primary document.

N
News / trade press

Journalism and trade coverage. Useful for events and controversies; weak for performance figures — the log's most-repeated caveat is a trade-press KPI that traces to no primary source.

V
Vendor claim

The weakest tier, admitted as evidence of claims only, never outcomes. A vendor's accuracy figure proves the vendor makes that claim — nothing more, until independently audited.

Adversarial review

The grill: every thesis is attacked before it's published

Between research and publication, every candidate thesis goes through a recorded adversarial session — the grill. Attack angles are run systematically: the evidence attack (which load-bearing facts rest on a single or self-interested source?), counterexamples inside and outside the domain, the constraint attack (was the binding constraint actually removed, or removed only in theory?), the incumbent's rebuttal (the strongest case against the thesis, argued in good faith), survivorship and selection bias, and falsifiability (what observable events would prove this wrong?).

A thesis can survive, die, or mutate. All three outcomes are recorded. Arguments that survive graduate into the analysis; arguments that die are kept as the objection-handling record; mutations are stated in the published piece rather than smoothed over. The Fashion & Luxury thesis, for example, mutated three times under the grill — its scope narrowed from "factory to resale" to "at manufacture, with the resale loop as the unproven next leg" — and that narrowing is stated plainly in the report's Blindspots section, along with three named falsifiers on a public watch list. One grill attack that could not be resolved became a public Unsolved Problems entry instead of a buried caveat.

The point of all three mechanisms — the tiers, the gaps, the grill — is the same: a skeptical reader should be able to audit how a report was built, not just read its output.