Frequently Asked Questions About Howie Liu Agent-First Business Builder

22 answers covering everything from basics to advanced usage.

// Basics

What does 'Frontier Agent, Frontier Model' mean?

It means always pairing your agent with the current highest-capability model — like Opus 4.5+ or GPT-5 — for serious work. The models are already smart enough for nearly every white-collar task, and underpowered models are the most common reason agents disappoint. Only downgrade to a cheaper model after Rubric data confirms quality is preserved.

What is Founder Mode for an agent?

Founder Mode is the configuration where an agent acts as a founder rather than just a developer — researching business context end-to-end, validating market need, mapping competition, and identifying structural market dynamics before it builds anything. App building is a feature inside Founder Mode, not the goal itself. This step separates informed builds from wasted effort.

What is the difference between gen-one AI and a frontier agent?

Gen-one AI is augmentation of still-human-driven workflows — tab autocomplete in an IDE, single-turn ChatGPT questions where a human stays in every step. A frontier agent operates fully autonomously across multi-turn tasks with tool access, capable of work that would take a skilled human hours or days, with no human in the per-step loop. Most people are still in gen-one mode without realizing it.

What is an LLM-as-judge and why does it matter?

An LLM-as-judge is a separate model instance that fires after each agent run, scores the output against your pinned Rubric, and returns dimension-level scores. It matters because you can't manually review every output as your fleet grows. It's automated quality oversight — Management 101 applied to agents, replacing manual inspection with checks and balances.

// How To

How do I create a Skill for a recurring job?

Have the agent research how the task should be done — including studying real examples of your style or domain — then distill that into a named, saved Skill. Specify what platforms it targets, what output format it produces, how autonomously it operates, and what constraints it respects. Pin it to a dedicated agent. Treat the Skill as a living playbook, not a one-time prompt.

How do I give feedback to improve an agent's output?

Review the first real outputs and identify one or two specific failure modes — for example, 'too formal, not colloquial enough' or 'no data supporting the claims.' Feed that back directly in the thread and have the agent both regenerate drafts and update the Skill so the fix becomes permanent. Fixing only the current output leaves future runs unchanged.

How do I build a Rubric for scoring agent output?

Define 3-5 dimensions that constitute 'great' for that agent's role, then have the agent help you build the Rubric — either via UI or by prompting in the thread with something like 'Help me build a rubric to score great X-style content.' Pin it to the agent. From then on, an LLM-as-judge scores every run and produces a quality trend line you can track.

How do I turn an on-demand agent into an always-on one?

Use one of two options: scheduled runs, where you tell the agent in the thread to run daily at a set time and deliver output via email or Telegram; or Live Mode, where the agent polls continuously — for example every 30 minutes — and pushes new drafts or ideas as they emerge. Keep draft-and-review for content; use full autonomy only for low-stakes tasks.

How do I expand from one agent to a full fleet?

Once one agent is stable, build the next for a different role — content marketer, market researcher, competitive intelligence, lead enrichment, customer email responder, or deal flow analyst. Give each its own Skills, Rubric, run schedule, and deployment target. Manage them all from the Command Center overview and one-click deploy any into Slack so teammates interact with it as a virtual co-worker.

// Troubleshooting

Why is my agent producing mediocre output?

The most likely causes are one-shotting and abandoning, using an underpowered model, or skipping the research and feedback loop. V1 output is always about 50% of your quality bar — that's normal and expected. Give specific interactive feedback, update the Skill so the fix is permanent, and pair the agent with a frontier model. Coaching over multiple iterations is what unlocks full capability.

My agent's quality seems to be degrading over time. What's wrong?

You're likely skipping the Rubric and Self-Improvement Loop. Without an LLM-as-judge scoring every run, quality degrades invisibly. Also check whether your Skills have gone stale — Skills that aren't continuously updated decay in relevance. Run memory defrag periodically to consolidate duplicate memories, and review the agent's suggested Skill and prompt updates rather than ignoring them.

The agent isn't personalizing recommendations to my business. Why?

You probably haven't connected it to your actual context. Starting from a blank slate without linking your Gmail, Slack, Notion, or Granola means the agent has nothing to personalize from. Connect these data sources first, then let it read your notes and communications so it can suggest use cases genuinely tailored to your situation rather than generic ideas.

I feel like I'm not getting value from agents. What am I doing wrong?

You're likely experimenting sporadically — trying the product once a week for a few minutes. Daily committed practice of at least 30 minutes for 30-60-90 days is the threshold for reaching top 1% proficiency. It's impossible to grasp what's now buildable without hands-on, ambitious use. Hand the agent a task that would take a skilled human many hours and let it run autonomously.

// Comparisons

How does the Agent-First framework compare to no-code app builders?

No-code app builders and prototyping toys stop at the artifact. The Agent-First framework treats app building as a commoditized feature inside a broader agentic workflow — the agent first researches the business, validates the market, and maps competition as a founder before building. It then adds Rubrics, fleet management, and self-improvement, which prototyping tools lack. It embodies Low Floor, High Ceiling: intuitive to start, scalable to run a real business.

How does this compare to hiring human employees for these roles?

Each agent is mapped to a human-equivalent role, but produces expert-level output at a fraction of the labor cost and time. A deal analyst agent can drop pitch review from 2 hours to 10 minutes; a content agent generates and scores daily drafts. Context window limits make role-partitioned agents structurally inevitable — the same reason companies partition human roles. Use the Human Equivalent Time Cost Reframe to compare fairly.

How does full YOLO mode compare to draft-and-review?

Full YOLO mode has the agent take autonomous action — posting content, sending emails, booking meetings — without human review, appropriate only for low-stakes, reversible, well-understood tasks. Draft-and-review keeps a human checkpoint before anything ships and is the correct default for content and external communications. Choosing the wrong mode for high-stakes output risks reputation damage; choosing it for trivial tasks wastes your time.

// Advanced

What is the Door-to-Door vs. Internet parable?

It's Howie's analogy for the agent-first transition. In the early internet, one salesperson dabbled with SEM on weekends while still selling door-to-door; another stopped selling entirely to master internet distribution for months. Two years of discomfort yielded a multi-billion-dollar outcome for the committed one. The same inflection is happening with agents — sporadic experimentation produces nothing, but committed daily practice compounds into structural business leverage within six months.

What is memory defrag and when should I run it?

Memory defrag is a periodic maintenance operation that clusters accumulated agent memories by keyword and embedding similarity, identifies duplicates and related items, and lets you consolidate them. Run it periodically as your agent's memory store grows, to keep it coherent and performant. Combined with curating the agent's suggested memory and Skill updates, it prevents memory bloat from degrading performance over time.

What market size should I target for an agent-first business?

Target a medium-sized market — a TAM of roughly a few billion dollars. It's large enough to build a multi-hundred-million-dollar business on a double-digit market share, but small enough that massive incumbents aren't prioritizing it. Avoid micro-niches that are too small to matter and hundred-billion-dollar categories that are too competitive. This is the sweet spot for a solo or small-team venture.

How does the Self-Improvement Loop actually work?

After runs, agents accumulate memories and surface suggested Skill updates, system prompt changes, and new tool access recommendations based on what they observed. You curate these suggestions — accept, reject, or modify — rather than accepting blindly. Over time the agent becomes progressively more effective at its role with decreasing intervention, but only if you actively coach and curate rather than one-shotting and abandoning.

What are the two named paths to building a valuable AI business?

Howie names two go-to-market paths: PLG (Product-Led Growth), where the product itself drives organic adoption through use without top-down sales; and the Palantir-style top-down enterprise check model, where you land large enterprise contracts through direct sales. Both are viable for agent-first businesses, and your choice shapes how you design the product's floor and ceiling and how you deploy agents into teams.

When is it worth switching an agent to a cheaper model?

Switch only after a Rubric is established and you have quality trend data. Test dropping from a frontier model like Opus to a mid-tier model like Sonnet on that agent's routine runs. If the Rubric score doesn't decline meaningfully, lock in the cheaper model — achieving up to 5x cost reduction. Keep frontier models reserved for high-stakes or complex tasks.