# The 1,000 questions to answer before building an AGI business agent

By Tom Carter, AlexDLY. Published October 8, 2026. Canonical page: https://www.alexdly.ai/blog/agi-business-agent-questions/

Before anyone writes a prompt, the best business builders interview the founder. They want the job, the data, the decision rights, the risk, and the money on the table. This is that interview, written out: 1,000 questions in 20 sections, in the order a serious build conversation runs.

## Contents

1. [Vision, ambition, and why now](#vision) (questions 1 to 50)
2. [What AGI means for your business](#agi-definition) (questions 51 to 100)
3. [The customer and the job to be done](#customer) (questions 101 to 150)
4. [Problem selection and the first wedge](#wedge) (questions 151 to 200)
5. [Business model, pricing, and unit economics](#economics) (questions 201 to 250)
6. [Market, competition, and moat](#market) (questions 251 to 300)
7. [Workflows and process mapping](#workflows) (questions 301 to 350)
8. [Data, knowledge, and organizational memory](#memory) (questions 351 to 400)
9. [Integrations, tools, and systems of record](#integrations) (questions 401 to 450)
10. [Autonomy, delegation, and human in the loop](#autonomy) (questions 451 to 500)
11. [Trust, verification, and evidence](#trust) (questions 501 to 550)
12. [Safety, security, compliance, and governance](#safety) (questions 551 to 600)
13. [Agent architecture and model strategy](#architecture) (questions 601 to 650)
14. [Evaluation, measurement, and quality](#evals) (questions 651 to 700)
15. [Reliability, operations, and cost](#operations) (questions 701 to 750)
16. [Personality, voice, and user experience](#experience) (questions 751 to 800)
17. [Go to market, distribution, and sales](#gtm) (questions 801 to 850)
18. [Onboarding, adoption, and change management](#adoption) (questions 851 to 900)
19. [Team, capital, and founder readiness](#team) (questions 901 to 950)
20. [Learning loops, scale, and the long game toward AGI](#long-game) (questions 951 to 1000)

<a id="vision"></a>

## 1. Vision, ambition, and why now

_Before building anything, the builder needs to know how big the founder wants to go and why this moment makes it possible. The answers set the scope, the risk tolerance, and every tradeoff that follows._

1. If this agent works exactly as you imagine, what does your company look like in 36 months in headcount, revenue, and customers?
2. Why are you the person to build this, and what do you know about this market that a well-funded competitor does not?
3. Why now: which specific change in model capability, cost, or buyer behavior in the last 18 months makes this possible today?
4. Which incumbent loses the most if you succeed, and how do you expect them to respond in the first year?
5. How big does this need to get before you would call it a success, and is that number in revenue, users, or enterprise value?
6. Are you building a product company, a services company with software, or an AI labor company, and which one do you want to be in five years?
7. If foundation model providers ship a native agent that does 70 percent of this next year, what is left that only you own?
8. Which sentence would you want a customer to say about your agent after 90 days of using it?
9. How much of your own time each week do you plan to spend on product, sales, and fundraising for the next six months?
10. Who is the first customer you can name today who would pay for this agent, and have you asked them?
11. Describe the moment you decided to build this: what happened, and who was in the room?
12. Where do you refuse to go, meaning which markets, use cases, or customers will you turn down even if they pay?
13. How long are you personally prepared to work on this if traction is slower than planned, in months?
14. Which existing company is closest to what you want to become, and where exactly does your path diverge from theirs?
15. When you say 'autonomous,' how many decisions per day do you expect the agent to make without a human in the loop by year two?
16. Is your ambition to replace a role, augment a team, or create a category of work that does not exist yet?
17. Which bet in your plan would you lose sleep over if it turned out wrong, and how would you find out early?
18. How will the world be different if you win, stated as one measurable outcome rather than a mission statement?
19. Should this company raise venture capital, bootstrap, or fund itself through services, and what does each path cost you?
20. Who has already tried this and failed, and what specifically will you do differently from them?
21. Which three assumptions must be true for this to be a billion dollar business, and how confident are you in each?
22. How fast is the underlying model cost curve dropping in your view, and does your plan get better or worse as it falls?
23. Which part of your vision would you cut first if you only had six months of runway left?
24. Where will your proprietary data advantage come from, and how many customers do you need before it compounds?
25. How would you explain this company to your grandmother in two sentences without using the word AI?
26. If you had to pick between owning the customer relationship and owning the best model routing, which would you keep?
27. Which regulatory or cultural shift could make your vision illegal, unwelcome, or obsolete within five years?
28. Who on your current team could run this without you for a month, and what would break first?
29. Are you building for the first 100 customers or the first 10,000, and which design choices change based on that answer?
30. How will you know in 12 months that you should pivot, and what signal would you accept as the verdict?
31. Which adjacent markets open up once the first agent is trusted, and in what order would you enter them?
32. What is the strongest argument against this company existing, as a skeptical investor would put it?
33. When a customer compares you to hiring one more employee, why does your agent win that comparison?
34. How much of your vision depends on models getting smarter versus your own engineering, workflow design, and data?
35. Which founders or operators would you call for advice on this, and what have they already told you?
36. Do you want this agent to have a name, personality, and brand presence, or to be invisible infrastructure, and why?
37. How public will you be about the errors the agent makes, and does that fit the kind of trust you want to build?
38. Which time horizon are you optimizing for right now: next quarter's revenue, next year's raise, or a decade-long platform?
39. Picture the version of this company that wins but you hate running; what does it look like, and how do you avoid it?
40. Is there an exit you would accept in three years, and at what price would you sell?
41. Which customer would make the best case study in a year, and why would their story convince the next thousand buyers?
42. How does your agent change the balance of power between small businesses and enterprises in your market?
43. Would you still build this if a lab released AGI tomorrow and any company could rent it for cents per hour?
44. Which metric will you report to your board every month as the single proof that the vision is working?
45. Where does your personal story give you distribution that a stranger could not buy?
46. How do you want employees of your customers to feel about the agent: relieved, threatened, or indifferent?
47. Which compounding advantage do you get from shipping first, and how long does that head start last?
48. Before writing any code, which result from ten customer conversations would you need to see to commit fully?
49. How much revenue per employee do you expect your own company to reach by using the agent internally?
50. Which belief about the future of work do you hold that most smart people in your industry would dispute?

<a id="agi-definition"></a>

## 2. What AGI means for your business

_AGI is a vague word until it becomes a list of tasks, limits, and tests. The builder pins down exactly what the agent must do, what it must never do, and how anyone will know it is finished._

51. When you say AGI for your business, which specific tasks must the agent perform end to end without a human touching them?
52. Where is the line between a smart workflow and general intelligence in your business, and which side does version one sit on?
53. How many distinct job functions should the agent handle by month 12: one, five, or every role in the company?
54. Which decisions is the agent allowed to make with money, and what is the dollar ceiling per action before it must ask?
55. If the agent faces a task it has never seen, should it attempt it, research it, ask for help, or refuse?
56. How will you measure generality, meaning the share of new requests it completes correctly without new code or prompts?
57. Which tools, systems, and accounts must the agent operate on day one, listed by name?
58. Should the agent learn from every interaction automatically, or should learning require human review before it changes behavior?
59. What does 'done' look like for the first release, stated as a test a stranger could run and score?
60. Which human skills are you explicitly not trying to replicate, such as negotiation, empathy, or creative direction?
61. How long should the agent be able to work on a single goal without supervision: minutes, hours, days, or weeks?
62. When two goals conflict, such as speed versus cost, who sets the priority order and where is it written down?
63. Which memory must persist across sessions: customer history, company policies, past mistakes, or all of them?
64. How does the agent know what it does not know, and how do you verify its confidence scores are calibrated?
65. Should it plan multi-step work on its own, or execute plans that a human approves first?
66. Can the agent spawn and manage other agents, and if so, who is accountable when a sub-agent fails?
67. Which outputs must be backed by evidence, like a receipt, log, or link, before the agent can claim success?
68. How should the agent behave when an external system is down, slow, or returns data it cannot trust?
69. Which level of autonomy do you want at launch on a scale from suggest only to fully autonomous, and at month six?
70. Is the agent a single persona the customer talks to, or a team of specialists behind one interface?
71. What benchmark from your own operations would convince you the agent performs at the level of your best employee?
72. Should the agent handle voice, email, chat, documents, images, and video, or which subset matters first?
73. How much context about the business must the agent absorb before it is useful, and how long should onboarding take?
74. Who defines the agent's values and red lines, and how are those enforced in code rather than in a prompt?
75. Which failure is worse for your customers: the agent doing nothing when it should act, or acting when it should wait?
76. How will you test the agent against adversarial inputs, such as a customer trying to talk it into a refund?
77. Does the agent need to explain its reasoning to users, and at what level of detail does that become noise?
78. Which model or models sit underneath, and how locked in are you if one provider changes price or policy?
79. Should the agent improve the business proactively, spotting opportunities nobody asked about, or only respond to requests?
80. How do you plan to separate what the agent reasons about from what it is permitted to touch?
81. When the agent makes a mistake that costs a customer money, what is the recovery playbook, step by step?
82. Which parts of the agent's capability come from the model, which from your tools, and which from proprietary data?
83. How do you prevent the agent from claiming work is complete when it only drafted or simulated it?
84. Can a non-technical owner teach the agent a new skill in under ten minutes, and how?
85. Which regulated actions, like hiring decisions, credit, health, or legal advice, are permanently out of scope?
86. How should the agent's behavior differ between a solo founder and a 200-person company using the same product?
87. Would you accept a less capable agent that is 99 percent reliable over a more capable one that is 85 percent reliable?
88. How often will you re-evaluate the agent against your capability bar, and who owns that evaluation suite?
89. Which data sources must the agent read in real time versus nightly, and what latency is acceptable for each?
90. Should users be able to see every action the agent took yesterday, and in what format?
91. Where should the agent hand off to a human, and how does that human pick up with full context in under a minute?
92. Does your definition of AGI include the agent setting its own goals, or only pursuing goals a human assigns?
93. How many concurrent tasks should a single agent instance run before quality degrades, and have you measured it?
94. Which competitor agent today comes closest to your capability bar, and where precisely does it fall short?
95. Who signs off that a new capability is ready for customers, and what evidence do they review?
96. How will the agent handle ambiguous instructions: ask a clarifying question, pick the most likely meaning, or offer options?
97. Is cross-customer learning allowed, and how do you keep one tenant's knowledge from leaking into another's answers?
98. Which three tasks, if the agent nailed them flawlessly, would make customers forgive weakness everywhere else?
99. When the next model generation ships, how quickly can you swap it in and prove nothing regressed?
100. Which capability are you tempted to promise in marketing that the agent cannot yet deliver reliably?

<a id="customer"></a>

## 3. The customer and the job to be done

_An agent is only as good as its fit with the person it serves and the person who pays. The builder needs names, numbers, and buying context, not a persona slide._

101. Who exactly is your first buyer, by title, company size, industry, and annual revenue?
102. Is the person who uses the agent the same person who pays for it, and if not, how do you win both?
103. Which task does your customer dread most each week, and how many hours does it take them today?
104. How does your customer solve this problem right now, and what do they spend on that solution per month?
105. Walk me through the last time a customer felt this pain: what happened, what did it cost, and who noticed?
106. How many businesses fit your ideal customer profile in your first geography, counted rather than estimated?
107. Which trigger event makes a customer start shopping for a solution like yours this month instead of next year?
108. Who inside the customer's company will resist the agent, and what do they lose if it succeeds?
109. How technically capable is your buyer, and can they connect their own systems without a call with you?
110. Which existing tools does your customer refuse to give up, and must the agent work inside them?
111. What does your customer believe about AI today, and how many bad experiences have they already had?
112. How does your buyer make purchase decisions: alone on a credit card, with a partner, or through procurement?
113. Which proof does your customer need before handing an agent access to their email, bank, or CRM?
114. Where does your customer spend time learning about new tools: podcasts, peers, LinkedIn, trade groups, or search?
115. How many customer interviews have you done in the last 30 days, and which answer surprised you most?
116. Which customer segment loves you most, and which one churns fastest, based on actual data?
117. Can you name ten customers who would be very disappointed if your agent disappeared tomorrow?
118. How does your customer measure their own success, and does the agent move that number directly?
119. Which jobs is your customer hiring the agent for that they would never admit to a vendor, like covering for an underperformer?
120. Who does your buyer report to, and what does that person need to see to keep renewing?
121. How sensitive is your customer's data, and which compliance requirements, like HIPAA or SOC 2, will they ask about?
122. Does your customer want to watch the agent work, or would they rather never think about it again?
123. Which language, tone, and level of detail does your customer expect in a status update from the agent?
124. How seasonal is your customer's business, and when in the year do they have budget and attention?
125. Should you serve owner-operators or managers inside larger firms first, and how does the product change for each?
126. Which competitor has your customer already tried, and why did they stop using it?
127. How long can your customer wait for value before they give up: one day, one week, or one month?
128. Who would your customer call first if the agent made a public mistake with one of their clients?
129. Which outcome would make your customer tell three peers about you without being asked?
130. How do your customers describe this problem in their own words, and are you using those words in your marketing?
131. Are you solving a painkiller problem or a vitamin problem for this buyer, and which evidence supports your answer?
132. Which customers should you fire or refuse because they will consume support without ever reaching value?
133. How does the customer's team change roles once the agent handles the work they used to do?
134. Where does your customer's data live today, and how messy is it on a scale you can quantify?
135. Which customer would pay double your price, and what is different about their situation?
136. How do you find customers who are already searching for this, versus educating customers who are not?
137. When your customer evaluates you, which three alternatives sit in the same spreadsheet column?
138. Is your buyer optimizing for saving money, saving time, making more revenue, or reducing risk, ranked in order?
139. How do you plan to keep talking to customers every week once you have hundreds of them?
140. Which part of the customer's workflow is emotionally loaded, where an agent's error would feel personal?
141. Who is the economic buyer at a 50-person company versus a 5-person company in your market?
142. How do your customers currently onboard a new human employee, and can the agent's onboarding mirror that?
143. Which customer requests have you said no to, and do those patterns point to a better market?
144. Can your customer articulate the problem before you explain it, or do you have to create the awareness?
145. What does a successful first week look like from your customer's perspective, day by day?
146. How many of your target customers have someone whose full-time job is the work your agent replaces?
147. Which channel brings in the customers with the highest retention, and why do you think that is?
148. Would your customer prefer a monthly report on what the agent did, or a live feed they can check anytime?
149. How does trust build with this customer over time, and which early win accelerates it most?
150. If you could only keep one customer segment, which one would you choose, and what would you give up?

<a id="wedge"></a>

## 4. Problem selection and the first wedge

_Agents that try to do everything on day one prove nothing. The builder forces a single, narrow workflow that shows value fast and earns the right to expand._

151. Which single workflow will the agent own first, and why that one over the other ten you could choose?
152. How narrow can you make the first use case while still delivering a result the customer would pay for?
153. Can the first workflow show measurable value inside the first 24 hours of a customer connecting their systems?
154. Which workflow has the highest frequency multiplied by pain, and have you scored the candidates side by side?
155. How many steps does the first workflow have today when a human does it, and which steps are judgment calls?
156. Which workflow is easy for the agent but hard for a human, giving you an unfair advantage on day one?
157. Where is the data for the first workflow already structured and accessible through an API?
158. How will you prove the agent completed the workflow correctly, with evidence the customer can check themselves?
159. Should the wedge be a task the customer currently ignores, or one they already pay someone to do?
160. Which workflow, once trusted, naturally pulls the agent into the next three adjacent jobs?
161. How long would it take a competitor to copy your first workflow, and what makes it hard to replicate?
162. Which tasks will you deliberately leave out of version one, even though customers will ask for them?
163. Can you run the first workflow manually, with humans behind the scenes, to validate demand before automating it?
164. How many customers must complete the first workflow successfully before you build the second one?
165. Which error rate in the first workflow is acceptable to customers, and how did you arrive at that number?
166. Does the wedge produce an artifact the customer can show their boss, like a report, a booked meeting, or a closed invoice?
167. Where in the workflow does the customer need to approve something, and can you reduce that to one tap?
168. Which integration is the critical path for the wedge, and what happens if that vendor changes their API?
169. How quickly can a new customer go from sign-up to the first completed workflow without talking to you?
170. Which version of the wedge would you be embarrassed to ship, and is that the version you should ship first?
171. If the wedge fails to retain users after 30 days, which metric will tell you whether to fix it or replace it?
172. How does the first workflow generate data that makes the agent better at the second workflow?
173. Which job in the customer's week happens every single day, so the agent earns trust through repetition?
174. Are you starting where the money is, such as revenue workflows, or where the risk is lowest, such as internal reporting?
175. How will you demo the wedge in under two minutes to someone who has never seen the product?
176. Which assumption about the wedge are you testing this week, and what result would make you abandon it?
177. Who in the customer's company owns the outcome of the first workflow, and do they feel the pain personally?
178. How do you avoid building a generic platform before a single workflow works flawlessly for paying customers?
179. Which customer requests are signals to expand the wedge, and which are distractions to politely decline?
180. Should the wedge be vertical, deep in one industry, or horizontal, one function across many industries?
181. How many hours per week does the first workflow save, and have you measured it with a stopwatch rather than a survey?
182. Which edge cases in the first workflow cause 80 percent of human escalations today?
183. Can you name the five design partners who will run the wedge first, and what did each commit to?
184. Will the wedge work with the customer's dirty, incomplete data, or only with the clean data in your demo?
185. Which success threshold, in conversion, accuracy, or time saved, triggers a full launch of the wedge?
186. How does pricing for the wedge set expectations for every workflow you add later?
187. When the wedge succeeds, how will the agent ask the customer for permission to take on more work?
188. Which workflow did you consider and reject as the wedge, and what would make you reconsider it?
189. How do you keep the wedge narrow when an early large customer offers money to build something broader?
190. Where does the wedge collide with an incumbent tool the customer already pays for, and do you integrate or replace?
191. Which part of the first workflow should stay human permanently because customers value the human touch there?
192. How many weeks will it take to get the wedge from prototype to a paying customer using it unsupervised?
193. Does the first workflow carry real consequences if it fails, and is that a feature for trust or a risk for adoption?
194. Which leading indicator in the first week predicts that a customer will still use the wedge in month three?
195. How will you capture every failure in the wedge so it becomes a test case rather than a support ticket?
196. Is the wedge something a customer would buy on its own, even if you never built anything else?
197. Which workflow would you pick if you had to show revenue impact to a skeptical CFO within 30 days?
198. How much custom setup does each new customer need for the wedge, and how do you drive that toward zero?
199. Which proof from the wedge goes into your sales deck, and how many customers must it represent to be credible?
200. When should you stop polishing the wedge and start expanding, and who makes that call?

<a id="economics"></a>

## 5. Business model, pricing, and unit economics

_An agent that runs on paid compute can lose money on every task it completes. The builder checks that value capture, cost to serve, and ROI proof add up before scale makes the math worse._

201. How will you charge: per seat, per task, per outcome, flat subscription, or a hybrid, and why?
202. What does it cost you in model tokens, infrastructure, and support to serve one customer for one month?
203. Which gross margin are you targeting at scale, and what is it today with real usage?
204. How much value does the agent create per customer per month in dollars, and what share of that do you capture?
205. If model costs drop 80 percent, do you pass the savings on, keep the margin, or expand the product?
206. Which customers are unprofitable at today's pricing because they use far more compute than average?
207. How long is your payback period on customer acquisition cost, in months, with your current numbers?
208. Should customers bring their own model API keys, and how does that change your margin and your liability?
209. How will you prove ROI to a customer in their own numbers, not a generic calculator on your website?
210. Which pricing tier will most customers land on, and is that the tier you make the most money from?
211. How do you cap compute spend when an agent loops or a customer triggers a runaway task?
212. Where does your price sit relative to the salary of the person whose work the agent performs?
213. Which costs scale with customers, which with usage, and which are fixed, in a breakdown you can show me?
214. How much does it cost to acquire a customer through each channel you use today?
215. Can you charge for outcomes like booked meetings or collected invoices, and how would you verify them without disputes?
216. Which expansion revenue path exists once a customer trusts the first workflow, and how large is it per account?
217. How do you price for a one-person business and a 500-person company without one subsidizing the other?
218. Which annual contract value would justify a sales team instead of self-serve, and when do you cross that line?
219. How much human labor sits behind each customer today, including onboarding and support hours?
220. Do you route tasks to the cheapest capable model, and how much does that save versus defaulting to the best model?
221. Which churn rate does your model assume, and how does lifetime value change if it doubles?
222. How will you handle customers who want a free trial of an agent that costs you real money every time it runs?
223. Where is your break-even point in customers and monthly revenue, and when do you expect to reach it?
224. Which price would make your best customer hesitate, and have you tested anything near it?
225. How sensitive is your margin to a single model provider raising prices by 30 percent?
226. Should you offer services, setup, or training as paid add-ons, and do they help or hurt the software multiple?
227. How do you reflect compute spend in your pricing when one task costs a cent and another costs five dollars?
228. Which metric will investors use to value you, and does your pricing make that metric look strong?
229. How much cash do you need to reach default alive, and which assumptions sit behind that number?
230. When a customer asks for a discount, which concession do you trade instead of lowering the price?
231. Which revenue would you lose if you refused to serve customers under a certain size?
232. How will you show customers exactly what they paid for, task by task, so the invoice never surprises them?
233. Does your pricing reward customers for using the agent more, or penalize them in a way that limits adoption?
234. Which partnerships, resellers, or marketplaces could lower your acquisition cost, and what share do they take?
235. How often will you revisit pricing, and which data will trigger a change?
236. What net revenue retention do you need for the business to compound, and how close are you today?
237. Can you explain your unit economics on one page that a CFO would sign off on without questions?
238. Which costs are you hiding from yourself today, like founder time on support or unpaid pilot work?
239. How do you price the agent's mistakes, including credits, refunds, or guarantees, and who absorbs that cost?
240. Which part of the stack should you build versus rent, given the cost and margin impact of each?
241. How does your cost to serve change as customers connect more tools and give the agent more context?
242. Would you rather have 100 customers at $1,000 a month or 1,000 customers at $100, and why?
243. Which pilot or proof-of-concept terms convert to paid contracts, and which ones just burn your time?
244. How will you price for agencies or consultants who deploy the agent across many of their own clients?
245. Where does the money come from in year one: subscriptions, services, or something else, by percentage?
246. How do you prevent caching, batching, and context size decisions from quietly eroding margins at scale?
247. Which guarantee could you offer, like money back if the agent does not save ten hours, without going broke?
248. How much revenue does a customer need to generate before your agent pays for itself, and how fast do they get there?
249. When usage spikes during a customer's busy season, does your pricing capture the upside or just absorb the cost?
250. Which number in your financial model are you least confident in, and what would it take to replace the guess with data?

<a id="market"></a>

## 6. Market, competition, and moat

_An agent nobody needs, or one a frontier lab gives away next quarter, is a bad build no matter how well it runs. The builder wants to know who pays, what they replace, and what still protects you when models get cheaper and smarter._

251. Who is your customer paying today to do the work this agent will do, and how much per month?
252. If OpenAI, Anthropic, or Google ships this exact agent as a free feature next quarter, why does your customer still pay you?
253. Which incumbent in your category already has the customer data, the distribution, and the budget to copy you in six months?
254. How many of your last 20 lost deals went to doing nothing, and what did the buyer say instead?
255. Which three alternatives does a buyer evaluate before you, and what is the single reason each one loses?
256. Where does your moat come from in year three: proprietary data, workflow lock-in, distribution, or brand, and what evidence supports it today?
257. How much historical context would a customer lose if they switched to a competitor tomorrow?
258. Which part of your product becomes worthless when the underlying model gets twice as smart and half as expensive?
259. Who inside the buyer's company signs the contract, and whose budget line does the money come from?
260. Is this agent replacing headcount, software spend, or agency fees, and which of those budgets is easiest to raise?
261. How do you price: per seat, per task, per outcome, or flat fee, and how does that survive agents doing the work of seats?
262. Which competitor has the best demo right now, and what does their demo hide that your customers would discover in week two?
263. How big is the slice of the market that will trust an autonomous agent with real writes this year, in number of companies?
264. Can you name ten companies who would pay for this today, and how many have already said yes in writing?
265. Why are you the team that wins this, versus a vertical SaaS vendor adding an agent tab to its existing product?
266. Which horizontal platforms, like Salesforce Agentforce or Microsoft Copilot, already sit inside your buyer's stack and claim this job?
267. How long does a customer need to use the agent before switching away becomes genuinely painful?
268. Does your agent get measurably better per customer over time, and how would you prove that compounding to a skeptical investor?
269. Which open source agent frameworks could a technical buyer use to rebuild 80 percent of your product in a weekend?
270. How would you describe the job your customer is hiring this agent for, in their words rather than yours?
271. Who has tried to build this before and failed, and what killed them?
272. Where does your gross margin land after model inference, tool calls, and human review costs per customer?
273. How sensitive are your unit economics to a 3x jump in token prices or a provider rate limit change?
274. Which single distribution channel brings you customers cheaper than any competitor can buy them?
275. How do you win against a services firm that promises the same outcome with humans and a money back guarantee?
276. Are you selling to the founder, the operator, or the IT buyer, and which of them can kill the deal?
277. What proof point would make a Fortune 500 operations leader trust a startup's agent with their CRM?
278. Which regulation or compliance standard in your market acts as a barrier you can clear and competitors cannot?
279. How fast is the customer's alternative improving, for example, how much better did ChatGPT get at this job in the last six months?
280. If you had to pick one vertical to dominate first, which is it and why that one over the next best?
281. How many customers would churn if you raised prices 40 percent tomorrow, and how do you know?
282. Which integrations or data partnerships would be hardest for a well funded rival to replicate?
283. Is the market buying AI agents as a category yet, or are you still educating buyers about the problem?
284. How does your agent's value show up on the customer's P&L within the first 90 days?
285. Which customers should you refuse to sell to, because serving them would pull the product off course?
286. When a frontier lab offers enterprise agents with SOC 2 and indemnity bundled, what is your counter in one sentence?
287. Where does the network effect live, if anywhere, between one customer's usage and another customer's results?
288. How defensible is your benchmark or evaluation data, and could a competitor generate equivalent data synthetically?
289. Which analyst, community, or influencer shapes your buyer's shortlist, and do they know your name?
290. How much of your current traction comes from founder relationships that will not scale past 50 customers?
291. Would your best customer recommend you to a direct competitor of theirs, and what would they say?
292. How do you plan to handle a price war when model costs collapse and every competitor drops to cost?
293. Which feature do competitors copy from you fastest, and what does that tell you about where your real edge is?
294. How long is your sales cycle today, and which step stalls most deals?
295. Can you show the retention curve for customers who reached first autonomous action versus those who never did?
296. Should this agent be a product, a platform, or a managed service, and which answer does your cash runway support?
297. How would your positioning change if you could only use the word agent once on your homepage?
298. Which adjacent market becomes reachable once this agent works for your first segment, and how much bigger is it?
299. What would a 10x better solution look like, and is anyone already building it?
300. Under what market conditions would you shut this product down rather than keep funding it?

<a id="workflows"></a>

## 7. Workflows and process mapping

_An agent can only run work that someone can describe, step by step, including where it breaks. The builder maps the real process, not the org chart version, before writing a single prompt._

301. Walk me through the last time this work was done end to end: who touched it, in what order, and how long did each step take?
302. Which single workflow, if the agent ran it flawlessly for 30 days, would change your business the most?
303. How many times per week does this workflow run, and what is the dollar value of one run going well?
304. Where are the written SOPs for this process, and how far does real practice drift from them?
305. Which step today depends on one person's judgment that nobody has ever written down?
306. How many handoffs between people or teams happen in this workflow, and where do items get dropped?
307. List the top five exceptions that break the happy path: how often does each one occur?
308. Who has the decision right at each step, and does the agent inherit that right or merely advise?
309. When the workflow stalls today, how does anyone notice, and how long does it take?
310. Which steps are pure data movement, which require judgment, and which require a relationship?
311. How do you define done for this workflow in a way a machine could verify?
312. Which inputs arrive unstructured, like emails, PDFs, voice notes, or screenshots, and in what proportion?
313. If the agent completes step four but fails on step five, what state is the work left in?
314. Which parts of this process exist only because of a past mistake or audit finding?
315. How long does a new hire take to perform this workflow competently, and what do they get wrong first?
316. Should the agent own the whole workflow or a slice, and where exactly is the boundary?
317. Which steps have hard deadlines, SLAs, or customer promises attached, and what happens when they slip?
318. How do priorities get set when three urgent items land on the same desk at once?
319. Who are the downstream consumers of this workflow's output, and what format do they expect?
320. Can you show me five real examples of this work done well and five done badly?
321. How does the workflow differ across your locations, brands, or business units?
322. Which steps would you happily delete entirely rather than automate?
323. When a customer asks for something off menu, who decides yes or no today, and on what basis?
324. How do approvals currently happen: Slack thumbs up, email reply, signature, or verbal, and which of those leave a record?
325. Which recurring meetings exist only to move this workflow forward, and could their output be a report instead?
326. How much rework does this process generate today, as a percentage of total effort?
327. What triggers the workflow to start: a calendar date, an inbound event, a threshold, or a person asking?
328. Where does the work wait the longest between steps, and why?
329. Which steps need to happen in strict order, and which could run in parallel?
330. How should the agent behave when two SOPs contradict each other?
331. Who is accountable when the workflow produces a bad outcome, and will that change when an agent runs it?
332. How do seasonal peaks change the volume and shape of this work?
333. Which steps require physical presence or a phone call that software cannot replace?
334. When was this process last redesigned, and who has authority to change it now?
335. How do you measure cycle time, quality, and cost for this workflow today, if at all?
336. Which tasks do your best people quietly do that are not in any job description?
337. If the agent finished this work in two minutes instead of two days, which downstream process would break from the speed?
338. How should the agent sequence follow-ups across email, SMS, and calls for the same contact?
339. Which judgment calls do your operators make by pattern recognition, and could they articulate the pattern if asked?
340. Are there legal or contractual steps, like disclosures or notices, that must happen in a specific wording?
341. Who should the agent hand off to when a workflow crosses from sales into delivery or finance?
342. How many distinct workflows do you want automated in year one, ranked by value and risk?
343. Which steps generate evidence that the work happened, and which steps leave no trace?
344. When your team ignores the process to get something done faster, what are they working around?
345. What is the cost of a single error in this workflow, in money, reputation, or legal exposure?
346. How should the agent handle a task that is half done by a human when it picks it up?
347. Which customer touchpoints in this workflow must feel personal, even if the agent drafts them?
348. Could you draw this workflow as a flowchart in 15 minutes, and if not, what makes it hard?
349. Which metric would your operations lead check first to know the agent is doing the job?
350. What would you need to see in a two week shadow run before letting the agent run this workflow live?

<a id="memory"></a>

## 8. Data, knowledge, and organizational memory

_An agent is only as good as what it knows, and most companies keep their truth in five places that disagree. The builder needs to know what the agent must remember, what it must forget, and whose data must never mix._

351. Which system is the single source of truth for customers, and which systems quietly disagree with it?
352. How many duplicate contacts, stale deals, or orphaned records sit in your CRM right now?
353. Where does your company's real knowledge live: docs, Slack threads, email, call recordings, or people's heads?
354. Which facts must the agent never forget, such as pricing commitments, customer promises, or legal holds?
355. Which data should the agent deliberately forget after a set period, and who sets those retention windows?
356. How do you want the agent to resolve two sources that state different prices for the same customer?
357. Who owns data quality for each source, and what happens when nobody fixes a known error?
358. How far back does useful history go, and when does old context start making answers worse?
359. If you run several businesses, how strictly must the agent keep their customers, data, and memory apart?
360. Which records contain personal, health, financial, or minor data that triggers special handling?
361. How should the agent cite where it learned a fact so a human can check it in one click?
362. When a customer asks you to delete their data, how will you purge it from the agent's memory and embeddings too?
363. Which lessons from past wins and losses should shape the agent's future decisions?
364. How do you capture decisions made on calls so the agent knows about them tomorrow?
365. Do you have a written glossary of your internal terms, product names, and acronyms the agent must understand?
366. How fresh does each data source need to be: real time, hourly, daily, or weekly?
367. Which documents are outdated but still get used, and how would the agent know they are stale?
368. Should the agent remember personal details about customers, like family names or preferences, and where is the line?
369. How will you test that the agent's memory of a customer is accurate, not plausibly invented?
370. Who can correct the agent's memory when it learns something wrong, and how fast does the fix propagate?
371. Which competitor, pricing, or market intelligence should the agent keep, and how do you stop it going stale?
372. How do you want memory to behave when an employee leaves: retained, reassigned, or reviewed?
373. Which data is too sensitive to send to any third party model provider, regardless of contract terms?
374. Are there regions or customers that require data residency in a specific country?
375. How much of your brand voice is documented, and how much exists only in examples?
376. Which metrics does your team argue about because different reports compute them differently?
377. How should the agent weigh a fact from last week against a contradicting fact from last year?
378. Can you export every core data source today, and in what format?
379. Which tribal knowledge would disappear if your two longest tenured employees quit tomorrow?
380. How should the agent separate what it observed, what it inferred, and what a human confirmed?
381. Who is allowed to see the agent's memory of a given customer, and does that mirror your existing permissions?
382. How do you want memory scoped: per user, per team, per business, or company wide?
383. Which outcomes, like closed deals or refunds, should be fed back to the agent as ground truth?
384. How large is your corpus of documents, emails, and transcripts in rough gigabytes or record counts?
385. If one customer's memory leaked into another tenant's answers, how would you detect it and how fast?
386. Which past customer complaints contain patterns the agent should learn from?
387. How do you label confidential deals, board matters, or HR issues so the agent stays out?
388. Should the agent keep a full audit history of every memory write, and for how long?
389. How do you onboard the agent's memory on day one: bulk import, guided interviews, or learning by watching?
390. Which fields in your systems are free text that people use inconsistently?
391. How will you know when the agent's memory has drifted from reality across hundreds of accounts?
392. Who decides which knowledge is canonical when your sales and operations teams describe the offer differently?
393. Which call recordings or meeting notes do you have consent to process, and which do you not?
394. How should the agent treat information it learns in a private DM versus a shared channel?
395. Does your pricing, policy, or product catalog change often enough that the agent needs versioned memory?
396. Which unanswered customer questions recur most, and is the answer written anywhere?
397. When the agent does not know something, should it ask, search, or say so, and in what order?
398. How do you want summaries of long customer histories formatted so a human can scan them in 30 seconds?
399. What is your plan for backups and point in time recovery if the agent's memory store is corrupted?
400. What evidence would convince you the agent knows your business better than a six month employee?

<a id="integrations"></a>

## 9. Integrations, tools, and systems of record

_The agent does real work only through the systems you already run, and every one of them has limits, quirks, and failure modes. The builder wants the access map and the breakage history before promising anything._

401. Can you list every system the agent must read from and write to, and who holds admin access to each one?
402. Which CRM do you run, how customized is it, and how many custom objects and fields matter?
403. Does each tool offer a real API, or will the agent need browser automation or exports?
404. Which systems must the agent only read, and which may it write to on day one?
405. How do you authenticate today: SSO, shared logins, personal API keys, or OAuth apps?
406. Who will own rotating and revoking the agent's credentials, and how quickly can you revoke them in an incident?
407. What are the rate limits on your critical APIs, and what happens to the workflow when the agent hits them?
408. When an API returns success but the record did not actually change, how would you catch it?
409. Which integration has broken most often in the last year, and how long did each outage last?
410. Should the agent send email from a shared inbox, a named person's account, or a dedicated sender domain?
411. How are SPF, DKIM, and DMARC set up on your sending domains, and has deliverability ever been a problem?
412. Which calendar system runs your bookings, and who is allowed to have meetings placed on their calendar automatically?
413. Does your ERP or accounting system allow write access through an API, and would your accountant permit it?
414. How do you want the agent to behave when a third party vendor changes its API without notice?
415. Which webhooks or event streams already exist, and which events are you currently polling for?
416. How much does each integration cost per call or per seat at the volume the agent will run?
417. Which systems are on plans that block API access, and what would upgrading cost?
418. Can the agent use a sandbox or test account for each system before touching production data?
419. How do you map one customer across your CRM, billing, support desk, and email when the IDs differ?
420. Which phone and SMS systems do you use, and do you have documented consent for automated messages?
421. How should the agent handle a partial failure where it updated the CRM but the email send failed?
422. Who at each vendor will you call when an integration fails at 2 a.m. on a Saturday?
423. Are any of your systems on premise or behind a VPN that the agent cannot reach from the cloud?
424. How will you prove to an auditor which actions the agent took in each system and under whose authority?
425. Which file stores, like Drive, SharePoint, or Dropbox, hold contracts and SOPs the agent needs?
426. How do you want idempotency handled so a retried action never creates duplicate invoices or emails?
427. Which integrations need per user OAuth because actions must appear under a specific employee's name?
428. Do your vendor terms of service allow an AI agent to act on your account and process the data?
429. How long can each system be unavailable before the agent should pause the workflow and alert someone?
430. Which reports or dashboards should the agent read rather than recomputing from raw data?
431. How should the agent tell an expired credential apart from a genuine permission denial?
432. Will you route through an integration platform like n8n, Zapier, or Make, or connect directly, and why?
433. Which fields in your CRM drive automations that the agent could trigger accidentally by writing to them?
434. How do you want payment systems like Stripe handled: read only, refunds allowed, or full charge capability?
435. Which systems store data in formats, such as scanned PDFs or images, that need extraction before use?
436. How many API calls per day do you estimate the agent will make at full volume, and does every vendor allow it?
437. Which integration should go live first because it unblocks the most value with the least risk?
438. How will you monitor integration health continuously instead of discovering failures from angry customers?
439. When a system of record and the agent disagree, which one wins, and who reconciles?
440. Which existing automations or Zaps will overlap with the agent, and which should be retired?
441. How should the agent handle attachments, signatures, and documents that need e-sign workflows?
442. Is there a staging copy of your data that mirrors production closely enough to test real edge cases?
443. Which systems log every change with a user stamp, and which overwrite silently?
444. Who approves adding a new tool to the agent's toolbox, and what review does that tool get first?
445. How should the agent behave when a page loads a cached error screen that looks like a valid response?
446. Which social, ad, and marketing platforms require app review before an agent can post or spend?
447. Can you list every place customer data flows out of your company today, including vendors and contractors?
448. How quickly can you disconnect the agent from every system at once if something goes wrong?
449. What scopes will you grant each OAuth app, and could narrower scopes still get the job done?
450. What would a single integration failure cost you in a day if nobody noticed it?

<a id="autonomy"></a>

## 10. Autonomy, delegation, and human in the loop

_The value of an agent is the work it does without you, and the risk is the same thing. The builder sets which actions run alone, which wait for a human, and how trust is earned and taken back._

451. Which actions may the agent take alone on day one, with no human approving them?
452. Which actions must always require human approval, no matter how good the agent's track record becomes?
453. How much money can the agent commit or refund without a human signature?
454. Who is the named human approver for each gate, and who backs them up when they are away?
455. How fast must approvals happen before the delay costs more than the risk the gate prevents?
456. Which actions are fully reversible, which are partially reversible, and which can never be undone?
457. Should the agent earn more autonomy over time, and what evidence threshold earns each new level?
458. When the agent is unsure, should it stop, ask, or proceed with its best guess and flag it?
459. How would you like escalations delivered: Slack, SMS, email, or an inbox, and with what urgency levels?
460. If an approver does not respond in four hours, should the request expire, escalate, or proceed?
461. Which customer-facing messages may the agent send without review, and which must a human read first?
462. Can the agent delegate tasks to your employees, and will they accept work assigned by software?
463. How should the agent follow up when a human it delegated to misses the deadline?
464. Who is liable when the agent takes an approved action that turns out to be wrong?
465. Which decisions do you personally want to keep even if the agent makes them better?
466. How will you audit a random sample of autonomous actions each week, and who does it?
467. What error rate would make you pull back autonomy, and over what time window?
468. Should the agent ever negotiate price, terms, or deadlines with a customer on its own?
469. How should the agent behave when two humans give it conflicting instructions?
470. Which approvals today are rubber stamps that could safely be removed?
471. How will you show approvers enough context to decide in under 30 seconds without opening another tool?
472. Do you want a kill switch per workflow, per integration, or one big button, and who can press it?
473. Should the agent be allowed to say no to a request from you, and in which situations?
474. How should the agent disclose to customers that they are dealing with an AI, and when?
475. Which actions should be batched for a single daily approval rather than approved one by one?
476. How will you prevent approval fatigue where humans click approve on everything without reading?
477. Could the agent take any action that affects an employee's pay, schedule, or performance record?
478. What should happen if the agent discovers fraud, a legal threat, or a safety issue mid workflow?
479. Who decides whether a new workflow starts in shadow mode, draft mode, or live mode?
480. How should the agent record why it chose an action so a reviewer can disagree with the reasoning?
481. Which time windows, like nights, weekends, or holidays, should restrict the agent's autonomy?
482. How should the agent report a task it attempted but could not confirm, so nobody mistakes an attempt for a result?
483. Should autonomy limits differ by customer tier, deal size, or region?
484. When a human edits an agent's draft before approving it, should the agent learn from that edit?
485. How should the agent escalate when the right human does not exist in your org chart yet?
486. Which actions should require two humans to approve, like large payments or contract changes?
487. Can the agent spawn sub-agents or call other agents, and who governs what they may do?
488. How will your team feel about an agent assigning them work, and what would make them trust it?
489. Is there any action the agent might take that would be legal but embarrassing on the front page?
490. How long should a pending approval sit before the agent re-checks whether the action is still valid?
491. Should the agent stop on its own when it notices something unusual, even without a defined rule?
492. Who reviews the approval rules themselves, and how often do they get revisited?
493. How should the agent behave during an outage of its approval channel?
494. Which outcomes will you measure to decide whether autonomy saved time or simply moved the work to reviewers?
495. Does your insurance carrier, board, or regulator need to know the agent acts on the company's behalf?
496. How should the agent hand a conversation back to a human mid thread without the customer noticing a seam?
497. Would you let the agent hire a contractor or buy a tool if it judged that the fastest path?
498. How do you want rollback to work when an autonomous action needs to be reversed hours later?
499. Which single autonomous action, if it went wrong, would you lose sleep over?
500. At what point in six months would you consider the agent a trusted colleague rather than a supervised tool?

<a id="trust"></a>

## 11. Trust, verification, and evidence

_An agent that cannot prove its work is a liability dressed up as a product. The builder needs to know what counts as proof before writing a single tool, because honest status has to be designed in, not bolted on._

501. When the agent says a task is done, which exact artifact, link, or record would you accept as proof?
502. How will you distinguish a simulated run from a live run in every screen, report, and notification your team sees?
503. Who on your team is allowed to mark an agent action as verified, and what evidence must they attach?
504. If the agent reports sending 40 emails, how do you confirm 40 actually left the mail server and reached inboxes?
505. Should the agent ever be allowed to report confidence higher than the evidence it can show for a claim?
506. How long must you retain the receipts for each agent action, and who needs to read them later?
507. Which three claims, if the agent got them wrong, would cost you a customer or a lawsuit?
508. Do you want the agent to say 'I don't know' more often, even if it makes the product feel weaker?
509. How should the agent report partial success, for example 7 of 10 CRM updates written and 3 rejected?
510. Where does the audit trail live, and can the agent itself edit or delete entries in it?
511. When a vendor API returns 200 but writes nothing, how should the agent detect and disclose that?
512. Which dashboards today show numbers you cannot trace back to a source row or API response?
513. How would you feel if a customer audited one week of agent work line by line tomorrow?
514. Should status labels be set by code from evidence, or can any human override them by hand?
515. If the agent learns from its own past outputs, how do you stop unverified claims becoming accepted facts?
516. How quickly must a false 'completed' status be caught and corrected before it damages trust?
517. Do your customers need to see the reasoning behind each action, or only the result and its receipt?
518. Which external systems can give you a durable ID or timestamp to prove an action happened?
519. How will the agent cite the source document when it states a fact about a customer or deal?
520. When two systems disagree, such as CRM and billing, which one does the agent treat as the truth?
521. Would you rather ship a feature labeled 'Built, not wired' or hold it back until it runs live?
522. Can you name a time an automation told you something worked when it did not, and what it cost?
523. How do you want stale data flagged, for example a pipeline value last synced nine days ago?
524. Should demo data ever appear in a live workspace, and if so, how is it watermarked?
525. Who signs off that a new integration is live, and which probe result do they require first?
526. How do you plan to measure trust itself: override rate, approval rate, or something customers report directly?
527. If the agent drafts a decision, how is the human approval recorded alongside the original draft?
528. How should the agent handle a cached error page that looks like a successful response?
529. Do you need tamper evident logs, such as hash chains, or is an append only table enough?
530. When the agent summarizes a meeting, how will a reader verify a quote against the transcript?
531. Which metrics in your pitch would survive a skeptical investor asking for the raw run ledger?
532. How should confidence scores decay when the underlying evidence ages or the source changes?
533. Is the agent allowed to retry silently, or must every retry appear in the record?
534. How do you want the agent to admit an error it made yesterday that nobody has noticed yet?
535. Should each handoff between agents carry its evidence forward, or start fresh with only a summary?
536. Who reads the audit trail weekly, and which decision changes based on what they find?
537. How do you prove to a regulator or enterprise buyer that a human approved a specific outbound message?
538. Can a customer export their full action history, with receipts, in a format their auditor accepts?
539. Which agent claims are currently unverifiable by design, and how will you label them honestly?
540. How do you prevent a heartbeat or 'running' indicator from showing activity that is not real work?
541. If the agent cannot reach a system to verify, should it block, warn, or proceed with a flag?
542. How will you reconcile agent reported revenue impact against your actual bank or Stripe deposits?
543. Do you want a single trust score per agent, or one per capability such as email, CRM, and billing?
544. How should a reviewer sample agent output: random, risk weighted, or every action above a dollar threshold?
545. When the agent quotes a number, must it show the query or filter that produced it?
546. How long can an agent action stay 'pending verification' before it escalates to a person?
547. Is there any situation where you would accept the agent's word without a receipt, and why?
548. How will customers tell the difference between your agent's opinion and a verified fact on screen?
549. Before you call a capability live, how many consecutive verified real runs do you require?
550. How do you keep marketing copy aligned with what the run ledger actually proves the agent did?

<a id="safety"></a>

## 12. Safety, security, compliance, and governance

_An autonomous agent holds keys to email, money, and customer data, so one bad instruction can become a breach or a lawsuit. The builder maps permissions, secrets, and legal exposure before granting the agent any power to act._

551. Which actions should the agent never take without a human click, regardless of how confident it is?
552. Where will API keys and OAuth tokens live, and can the model ever see their raw values?
553. How will you stop a malicious email or web page from instructing the agent to exfiltrate data?
554. Which categories of personal data will the agent touch, and which jurisdictions' privacy laws apply to them?
555. Does each model provider you plan to use train on your inputs, and have you signed their DPA?
556. Who can press the kill switch, how fast does it stop in-flight actions, and has anyone tested it?
557. How granular must permissions be: per tool, per record, per customer, or per dollar amount?
558. If the agent sends a defamatory or false message to a client, who carries the liability?
559. Are you in a regulated industry such as health, finance, or legal, and which rules bind agent output?
560. How will you keep one tenant's data from appearing in another tenant's prompts, memory, or retrieval results?
561. Should the agent run with the founder's credentials, or with scoped service accounts that have least privilege?
562. How do you plan to rotate a leaked secret, including re-encrypting everything protected by the old key?
563. Which outbound channels, like SMS or cold email, require consent records and opt-out handling before the agent sends?
564. How will you red team the agent before launch, and who outside your team will attack it?
565. Can the agent spend money, and if so, what hard daily cap applies and who sets it?
566. How will tool outputs be labeled as untrusted data so the model never treats them as instructions?
567. Do your enterprise prospects require SOC 2, ISO 27001, or a security questionnaire before a pilot?
568. Where are prompts and completions logged, and do those logs contain PII that needs redaction or retention limits?
569. Which data must stay in a specific region, and can your chosen model providers guarantee that?
570. How do you handle a customer request to delete their data from memory, logs, and any fine-tuned weights?
571. Should the agent be able to modify its own instructions, tools, or permission settings?
572. How will you scope OAuth grants so the agent gets read only access unless writes are justified?
573. Who reviews new tools before the agent can call them, and which checklist do they use?
574. How do you prevent the agent from scraping or using data in ways that violate a vendor's terms?
575. Are you comfortable with a model provider retaining prompts for 30 days for abuse monitoring, or do you need zero retention?
576. How should the agent behave when a user asks it to do something legal but against your policy?
577. Which employees can view agent transcripts, and how is that access logged and reviewed?
578. If an attacker compromises one connected SaaS account, what blast radius does the agent give them?
579. How do you plan to disclose to recipients that a message was drafted or sent by an AI agent?
580. Does your insurance cover errors made by an autonomous agent acting on behalf of clients?
581. How will you enforce separation between the founder's personal accounts and business accounts the agent uses?
582. Can the agent's sandbox reach the open internet, and which domains are on its allow list?
583. Which actions are irreversible, and how will the agent confirm intent before executing any of them?
584. How will you handle hiring, advertising, or lending decisions where algorithmic bias rules may apply?
585. Do customers bring their own model keys, and if so, who is responsible when their provider mishandles data?
586. How do you stop the agent from pasting confidential documents into a third party tool for convenience?
587. Who owns the governance policy for the agent, and how often is it reviewed against new incidents?
588. If a court subpoenas agent logs, what will you be able to produce, and what should you never store?
589. How will role based access ensure a junior staffer cannot direct the agent to read payroll data?
590. Should the agent refuse tasks for a customer whose account is past due or under a legal hold?
591. How will you version and approve changes to system prompts, given they effectively change policy?
592. Is there a written incident policy for when the agent leaks data, including customer notification timelines?
593. How do you verify that third party MCP servers or plugins are not quietly sending data elsewhere?
594. Which model providers are off limits for sensitive data, and how does the router enforce that rule?
595. Do your contracts with customers permit an AI agent to process their data and their customers' data?
596. How will you detect an agent session that has been hijacked mid conversation by injected content?
597. Should minors ever interact with the agent, and which age gating applies if so?
598. How do you audit that the agent respected a do not contact list across every channel?
599. What would a breach notification to your top ten customers say if it happened next month?
600. Before granting a new capability, who must sign off, and what written rollback plan is required?

<a id="architecture"></a>

## 13. Agent architecture and model strategy

_The shape of the system decides what the agent can do in year two, and many early choices are expensive to reverse. The builder tests whether the founder has a reason for each layer: orchestration, routing, tools, and memory._

601. Should this start as one agent with many tools, or several specialist agents with a coordinator, and why?
602. Which tasks need multi step planning, and which are better as fixed workflows with a model at one step?
603. How will you route between cheap fast models and expensive frontier models for each task type?
604. Do you want customers to bring their own model keys, and how does that change routing and margins?
605. Where does long term memory live: vector store, relational tables, a knowledge graph, or plain documents?
606. How will the agent decide what to remember, what to forget, and what requires human confirmation to store?
607. Which tools must be built in house, and which can you buy or wrap from existing APIs?
608. How will you handle a model provider deprecating the model your prompts are tuned for?
609. Should orchestration run on a durable workflow engine, a queue, or inside a long lived chat session?
610. How will the agent recover state if it crashes halfway through a ten step task?
611. How many tools can the agent see at once before selection accuracy drops, and how will you test that?
612. Is n8n, Temporal, LangGraph, or a custom loop the right runtime for your first production workflows?
613. How should context be assembled per turn so the model sees the right facts without blowing the token budget?
614. Which decisions belong to deterministic code rather than the model, even if the model could make them?
615. How will you represent each customer's business: a schema, a profile document, or learned embeddings?
616. Do agents communicate through shared state, direct messages, or a task board, and how are conflicts resolved?
617. How will tool definitions be versioned so an API change does not silently break the agent?
618. Would fine tuning help any task, or is retrieval plus good prompting enough for the next twelve months?
619. How will you structure prompts so policy, persona, tenant context, and task instructions do not collide?
620. Which integrations require write access on day one, and which can stay read only for the first quarter?
621. How should the agent schedule recurring work, and who owns the clock: the agent, cron, or the customer's calendar?
622. Will the agent act through APIs only, or also through browser and desktop automation when no API exists?
623. How will the router choose a fallback model that preserves the tool calling behavior your prompts depend on?
624. Should each tenant get an isolated agent instance, or share one runtime with strict scoping?
625. How will retrieval rank sources when the CRM, inbox, and documents all claim different facts?
626. When should the agent ask a clarifying question versus making a reasonable assumption and proceeding?
627. How will human approvals plug into the agent loop without blocking unrelated work behind them?
628. Which open source components will you depend on, and what happens if a maintainer abandons one?
629. How do you plan to give the agent structured outputs that downstream code can parse every time?
630. Does the agent need real time voice, or are text and asynchronous channels enough for your buyers?
631. How will you separate the agent's planning model from its execution model, if at all?
632. Where should business rules live so a non engineer can change them without a deploy?
633. How will the agent handle tasks that span days, like waiting for a prospect to reply?
634. Which parts of your stack would you rip out first if a competitor launched the same agent cheaper?
635. How does the agent learn from approved edits, and how do you stop it learning from rejected drafts?
636. Is MCP your integration standard, and how will you expose your own capabilities to other agents through it?
637. How will you keep the system prompt stable enough to cache while still carrying tenant specific rules?
638. Can you draw the data flow from a user request to a completed external action on one page?
639. How will multimodal inputs like screenshots, PDFs, and call recordings enter the agent's context?
640. Should the agent have one persona across every customer, or configurable roles per deployment?
641. How will you size the context window you actually need, given most tasks may need only 20K tokens?
642. Who owns the architecture decisions, and where are they written down so new engineers follow them?
643. What is your build versus buy line for evals, observability, and memory infrastructure?
644. How will the agent know which business, brand, or workspace it is acting for on each request?
645. Do you need an internal registry where every agent is declared, with owner, tools, and permissions, before it runs?
646. How will you keep the agent's model of the org chart current as people join and leave?
647. Which single capability, if architected wrong now, would be hardest to change in a year?
648. How will reasoning models' longer latency fit into workflows where users expect an answer in seconds?
649. Should the agent be able to spawn sub agents dynamically, and what limits stop runaway recursion?
650. How will you test a new model against your production prompts before switching any customer traffic?

<a id="evals"></a>

## 14. Evaluation, measurement, and quality

_Without a fixed set of real tasks and an honest score, every model swap or prompt edit is a guess. The builder wants to know how the founder will tell better from worse before customers find out first._

651. Which twenty real tasks from your business would you use as the first golden evaluation set?
652. How do you define success for one agent run in a way two reviewers would score identically?
653. Who on your team will grade outputs weekly, and how many hours can they actually commit?
654. Should an LLM judge score outputs, and how will you calibrate it against human graders?
655. Which single metric would tell you the agent is getting worse before a customer complains?
656. How will you collect failed runs from production and turn them into new eval cases?
657. Do public benchmarks matter to your buyers, or only performance on their own workflows?
658. How will you measure business outcomes, like booked meetings or collected invoices, rather than output quality alone?
659. How large must an eval set be before a two point score change is statistically meaningful to you?
660. Should evals run on every prompt change, every deploy, or nightly against production snapshots?
661. How will you score tasks with no single right answer, like a sales email or campaign brief?
662. Which failure is worse for you: the agent doing the wrong thing, or doing nothing when it should act?
663. How do you plan to weight errors by cost, so a wrong invoice counts more than a typo?
664. Can you get customers' permission to use their real runs as anonymized test cases?
665. How will you evaluate multi step tasks where the final output looks fine but an intermediate step was wrong?
666. What human edit rate on drafts would convince you the agent is ready for less supervision?
667. How will you track whether users accept, edit, or discard each output, and feed that back into scoring?
668. Which competitor or human baseline will you compare against, and how will you run that comparison fairly?
669. How do you stop the eval set from leaking into prompts or examples, inflating your scores?
670. Should each tenant have its own eval suite reflecting their voice, rules, and edge cases?
671. How will you measure tool selection accuracy separately from final answer quality?
672. When a new model scores higher overall but fails three critical cases, do you ship it?
673. How will you detect regressions in tone or brand voice that numeric scores might miss?
674. Who decides the pass threshold for each capability, and which evidence justified that number?
675. How often will you refresh golden tasks so they still reflect how customers actually work?
676. How do you score refusals: as safe successes, as failures, or by whether the refusal was correct?
677. Do you want a public scorecard customers can see, and which numbers are you willing to publish?
678. How will you evaluate memory recall, meaning whether the agent surfaces the right fact at the right moment?
679. How will you test adversarial inputs, like hostile customers or injected instructions, inside the eval suite?
680. Which latency and cost thresholds count as eval failures, even when the output is correct?
681. How long does one full eval run take today, and how fast must it be to fit your release cycle?
682. How will you rubric a strategic recommendation, where quality only shows up months later?
683. Can you name the last time a metric looked good while the underlying experience got worse?
684. How will you measure time saved per user in a way a skeptical CFO would accept?
685. Should eval results gate deploys automatically, or inform a human who makes the call?
686. How do you plan to evaluate the agent's honesty about its own uncertainty and limitations?
687. Which tasks will you evaluate with replayed real data versus synthetic scenarios, and why?
688. How will you version eval sets so a score from March is comparable to one from October?
689. What does a 70 out of 100 score mean in practice for a customer deciding whether to rely on it?
690. How will you sample production traffic for human review without reviewers burning out on volume?
691. Do you track inter rater agreement, and how do you resolve cases where graders disagree?
692. How will you evaluate consistency, meaning the same input produces acceptably similar output across runs?
693. Which customer segment's workflows are underrepresented in your evals today, and why?
694. How will you measure the quality of handoffs between agents, not just each agent alone?
695. Should you pay domain experts, such as accountants or recruiters, to grade outputs in their field?
696. How will you test that the agent follows each tenant's written policies, not just general good practice?
697. Who gets alerted when a nightly eval score drops, and what must they do within 24 hours?
698. How will you know if customers churn because of quality, rather than price or onboarding?
699. How do you plan to evaluate actions taken in external systems, where you cannot easily replay the result?
700. Which three eval results would you show an investor to prove the agent works on real business tasks?

<a id="operations"></a>

## 15. Reliability, operations, and cost

_A business agent that works in a demo but fails at 2 a.m. or burns cash per task is not a business. The builder pins down uptime, failure handling, observability, and unit cost before customers depend on it._

701. What uptime does a customer need from the agent, and what happens to their business during a one hour outage?
702. How will you fail over when a model provider returns errors, and which tasks may never fail over to a cheaper model?
703. Which traces, spans, and logs will let you reconstruct any agent run within five minutes?
704. Who is on call for agent incidents at 2 a.m., and what can they do without waking an engineer?
705. What is your target cost per completed task, and at which price point do the unit economics break?
706. How will you cap token spend per tenant so one runaway loop cannot burn a month's budget overnight?
707. How do you handle rate limits from vendors like Google, HubSpot, or GoHighLevel at peak volume?
708. Which latency budget applies to interactive chat versus background tasks, and how will you enforce each?
709. How will you know a scheduled job silently stopped running three days ago?
710. Can your database absorb heartbeat and telemetry writes at 100 times today's volume without filling the disk?
711. How will you roll back a bad prompt or model change in under ten minutes?
712. Do you have a status page, and will it report degraded agent quality, not just downtime?
713. How should the agent queue work when a dependency is down, and how long can work safely wait?
714. Which costs grow linearly with customers, and which grow with how much each customer uses the agent?
715. How will you attribute compute spend to each tenant, feature, and agent for pricing decisions?
716. How do you plan to run a postmortem after an agent incident, and who sees the write up?
717. How will you test failover paths regularly, rather than discovering they fail during a real outage?
718. Does prompt caching apply to your workloads, and how much would it cut your monthly model bill?
719. How will you deploy changes: canary tenants, feature flags, or all at once?
720. Which alerts would page someone, and which would only land in a daily digest?
721. How do you handle expired OAuth tokens across hundreds of customer integrations without manual support tickets?
722. How many concurrent agent runs can your current infrastructure handle before queues back up?
723. Should batch workloads move to discounted asynchronous model APIs, and how much delay can customers tolerate?
724. How will you detect a retry storm where the agent hammers a failing API in a loop?
725. Who owns the vendor relationships and billing for each model and SaaS provider you depend on?
726. How will database migrations run safely on both a fresh install and production without manual steps?
727. Which single point of failure in your stack would take every customer down at once?
728. How do you plan to back up memory, logs, and configuration, and when did you last test a restore?
729. Will you run the agent in one cloud region or several, and what drives that choice?
730. How will customers learn that a task failed, and how quickly after the failure?
731. What runbook exists for a provider billing failure, such as an API key on an account with zero credit?
732. How will you measure p95 latency per tool call so slow integrations are visible, not hidden in averages?
733. How much engineering time per week goes to keeping integrations alive, and is that sustainable?
734. Should heavy jobs like crawling or media generation run on separate workers from interactive requests?
735. How will you forecast compute spend for next quarter given uncertain customer adoption?
736. Do you have a dead letter queue for failed actions, and who reviews it each morning?
737. How will you support a customer who says the agent did something you cannot find in the logs?
738. Which operational metrics belong on the founder's weekly dashboard alongside revenue and churn?
739. How will you keep staging data realistic enough to catch production issues without copying customer PII?
740. How long can you tolerate degraded mode, such as read only operations, before customers churn?
741. Which vendors would hurt most if they changed pricing tomorrow, and how would you exit each one?
742. How will you handle timezone, daylight saving, and holiday edge cases in scheduled agent work?
743. Is there a written service level agreement you will offer, and which credits will you pay on misses?
744. How will you cap the number of tool calls or steps per run so loops terminate predictably?
745. Who can deploy to production, and which review gate applies before any push?
746. How do you plan to scale support when each new customer brings unique integrations and edge cases?
747. How will you measure the gross margin of each customer after model, infrastructure, and support costs?
748. How will you verify that your health checks exercise the full request path rather than a default response?
749. How will you warm up the system after a full outage so queued work does not overwhelm vendors?
750. Which operational task do you still do by hand weekly that should be automated before you scale?

<a id="experience"></a>

## 16. Personality, voice, and user experience

_People decide whether to trust an agent in the first few exchanges, long before they judge its results. The builder needs to know how the agent should sound, where it lives, and how it proves its work so the experience earns trust instead of spending it._

751. If Alex had to introduce itself to your newest employee in two sentences, what exactly would it say?
752. Which three adjectives describe the voice you want, and which three would make you cancel your own subscription?
753. Should the agent ever use humor with your customers, and if so, give me one line you would actually approve?
754. How should the agent's tone shift between a frustrated customer, a busy CEO, and a junior hire asking a basic question?
755. Do your users live in chat, a console, Slack, email, or their phone, and which one wins when they conflict?
756. Where does a user first meet the agent: a login screen, a text message, a Slack DM, or an email from it?
757. How many seconds can a user wait for a response before the experience feels broken in your market?
758. When the agent is working on a task for twenty minutes, what does the user see during that time?
759. How should the agent show its work: full reasoning, a short receipt, a link to evidence, or nothing unless asked?
760. Which decisions must the agent explain before acting, and which can it simply report after the fact?
761. How often may the agent interrupt a user unprompted each day before it becomes noise they mute?
762. Which five events are important enough that the agent should push a notification to someone's phone immediately?
763. Should the agent speak in first person as a named teammate, or as a neutral system, and why for your buyers?
764. How will the agent admit it does not know something or cannot do a task without eroding trust in everything else it says?
765. When the agent is uncertain, should it show a confidence score, a plain caveat, or ask a clarifying question first?
766. How many clarifying questions may the agent ask before a user feels interrogated rather than helped?
767. Which tasks should be fully conversational, and which need a form, table, or button because chat is too slow?
768. Does the console need to work for a non-technical owner on a phone at 6 a.m., or for an analyst on two monitors?
769. How will a user undo something the agent did, and how many clicks or words should that take?
770. Should the agent remember personal preferences like report format and meeting times, and who can see that memory?
771. How long should a typical agent reply be, in words, and how will you enforce that limit?
772. Which formatting choices, such as bullets, tables, bold text, or emojis, are banned in your brand voice?
773. Will the agent have a voice interface, and which tasks would someone actually do by talking instead of typing?
774. If the agent sends an email on a user's behalf, should recipients know it was written by an agent?
775. How should the agent behave when two users in the same company give it contradictory instructions?
776. Do you want a morning brief, an end-of-day summary, a weekly review, or none, and what goes in each?
777. How will the agent display money, deadlines, and risk so a skimming executive catches them in three seconds?
778. Which single screen would a power user keep open all day, and what are the four things on it?
779. How should the agent handle a user who is rude, abusive, or trying to jailbreak it in front of colleagues?
780. When the agent finishes a task, what proof does the user get: a link, a screenshot, a record ID, or a receipt?
781. How will you test whether the agent's personality survives a model swap from one provider to another?
782. Should the agent mirror each user's writing style, or keep one consistent company voice everywhere?
783. Which accessibility needs, such as screen readers, dyslexia, or low bandwidth, show up in your actual user base?
784. In which languages must the agent operate on day one, and who will judge its quality in each?
785. How does a new user discover what the agent can do without reading documentation or watching a tour?
786. Would your users prefer the agent propose a plan and wait, or act immediately and report back, for routine tasks?
787. Can you show me one real message from your best employee that captures the tone you want the agent to copy?
788. How should empty states look on day one, before the agent has any data about the user's business?
789. Where should approvals happen: in chat, in an inbox, by text reply, or in the tool where the work lives?
790. How will the agent signal the difference between simulated, drafted, and live actions so nobody confuses them?
791. Who on your team owns the agent's voice guidelines, and how often will they review real transcripts?
792. Which three user complaints about existing AI tools do you want your experience to make impossible?
793. Should the agent have a face, an avatar, or a name at all, and how did you test that choice with buyers?
794. How will mobile differ from desktop: a full mirror, an approvals-only app, or a text-message interface?
795. When the agent makes a mistake, what are the exact words it should use to own it and fix it?
796. How do you measure delight in this product, beyond thumbs up and thumbs down on individual replies?
797. Is there any situation where the agent should refuse to answer a user's question even though it technically could?
798. How much latency will you trade for quality, for example waiting ten seconds for a better answer instead of two?
799. Which part of your current interface do users ignore entirely, and why would an agent version be different?
800. If a user opened the product and saw only one sentence from the agent, what should that sentence be?

<a id="gtm"></a>

## 17. Go to market, distribution, and sales

_A brilliant agent that nobody buys is a research project. The builder needs to know exactly who pays, why they pay now, and how the next hundred customers will find you without the founder in every deal._

801. Who exactly is your ideal customer by title, company size, industry, and the trigger event that makes them buy this quarter?
802. Can you name the first ten customers you could close by Friday, and explain how you reach each one?
803. Which existing budget line does your agent replace: headcount, agency fees, software seats, or outsourced services?
804. What is your price, how did you arrive at it, and how many customers have paid it without a discount?
805. Will you sell per seat, per agent, per outcome, or per task completed, and how does that scale with value?
806. How many minutes into a demo does the buyer see the agent do real work on their own data?
807. Which channel produced your last three customers, and can it produce thirty more at the same cost?
808. How much does it cost you today to acquire one paying customer, including your own time at a real hourly rate?
809. Who is the economic buyer, who is the champion, and who is the blocker in a typical deal?
810. Which competitor do prospects mention most often, and what do you say when they bring it up?
811. Why would a buyer choose you over hiring a virtual assistant for $1,500 a month?
812. How will you explain what the agent does in one sentence to someone who has never used ChatGPT?
813. Does your pricing page show prices publicly, and what happens to conversion if you hide them?
814. Which partners already sell to your buyer, such as agencies, accountants, or consultants, and what is in it for them?
815. How will you show up when a buyer asks ChatGPT, Claude, or Perplexity for the best tool in your category?
816. Which ten search queries or AI prompts do you need to own, and what content answers each one today?
817. How long is your sales cycle from first call to signed contract, and where do deals stall most?
818. Is this product-led, sales-led, or services-led for the first year, and what evidence chose that?
819. How many outbound messages per week can you send without burning your domain or LinkedIn account?
820. Which case study, with real numbers and a named customer, will you have published within ninety days?
821. Should you offer a free trial, a paid pilot, or a done-for-you setup, and which converts best in your tests?
822. How will you handle procurement questions about security, data residency, and SOC 2 before you have certifications?
823. What is your annual contract value target, and how many customers at that price make the company default alive?
824. Which vertical will you dominate first, and what will you deliberately refuse to sell to until then?
825. How does a customer refer another customer, and what reward actually motivates your buyer type?
826. Who on your team runs demos when you are on vacation, and can they close without you?
827. Which objection kills the most deals today, and what proof neutralizes it?
828. How will you price usage spikes, such as an agent that suddenly runs ten thousand tasks in a month?
829. Are you selling the agent, the outcome, or the transformation, and which one does your website lead with?
830. How many qualified demos per month do you need to hit next quarter's revenue goal, given your close rate?
831. Which communities, newsletters, podcasts, or events does your buyer actually trust, and are you known there?
832. Would a listing in the HubSpot, Slack, or Shopify app marketplaces bring real buyers, or just installs that never activate?
833. How will you measure whether content marketing drives pipeline, not just traffic and likes?
834. When a prospect says 'we are building this internally,' how do you respond and what is that response based on?
835. Which discount policy will you hold firm on, and who has authority to break it?
836. How much of your revenue should come from consulting or setup fees versus recurring software in year one?
837. Do your landing pages have a call to action above the fold, and what does it ask the visitor to do?
838. Which free tool or audit could you offer that proves value before a buyer ever talks to sales?
839. How will you sell into a company where the owner loves AI but the staff is openly skeptical?
840. Where do your best leads come from geographically, and does that change how you sell or support?
841. Is there a channel partner who could bring you a hundred customers, and what would they need to see first?
842. How will you prevent the founder from being the only person who can sell the product?
843. Which metric on your sales dashboard do you check every morning, and what number makes you worried?
844. Should your annual plan include implementation, and what does the customer get on day one for the money?
845. How will you handle a large customer who asks for custom features as a condition of signing?
846. What is your plan when a well-funded competitor copies your positioning and outspends you on ads?
847. Can you show me the email or message that booked your most recent demo, word for word?
848. How will expansion revenue work: more agents, more departments, more locations, or higher tiers?
849. Which buyer personas have you tried and failed to sell to, and what did you learn from each loss?
850. If you had $10,000 for go to market next month, exactly where would each dollar go?

<a id="adoption"></a>

## 18. Onboarding, adoption, and change management

_Most AI deployments fail inside the customer, not inside the code. The builder asks these to find out how fast the agent proves value, how staff fears get handled, and what keeps customers from quietly churning._

851. How many minutes from signup until a new customer sees the agent complete one real task on their business?
852. Which single first task will you guide every new account to, and why that one?
853. Who inside the customer's company is responsible for adoption, and what happens when that person leaves?
854. How will you reassure staff who believe the agent is being installed to replace them?
855. Which employees should get access first: the enthusiasts, the managers, or the people with the most repetitive work?
856. How many training sessions does a typical customer need before using the agent without help, and who runs them?
857. Which data connections are mandatory on day one, and how many customers abandon setup at that step today?
858. How will you know within the first week that an account is heading toward churn?
859. Which daily habit should the agent anchor to, such as the morning inbox, the standup, or the end-of-day review?
860. How will you handle a customer whose data is so messy that the agent's first outputs are wrong?
861. Which permissions should start locked, and what milestones grant the agent more autonomy over time?
862. How will managers see which team members use the agent, without turning it into surveillance?
863. Does onboarding happen self-serve, on a call with your team, or through a partner, and what does each cost you?
864. Which in-product moments will you use to teach new capabilities without a separate training course?
865. How many weekly active users per account signals a healthy customer in your model?
866. When a customer says 'it did not work,' how do you diagnose whether it was the product, the data, or the setup?
867. How will the agent surface early wins so the customer's team notices and talks about them?
868. Which department inside a customer usually resists hardest, and what is your playbook for them?
869. Is there a written change management plan you hand to customers, and who has actually followed it?
870. What will you do for the customer who signed up, connected nothing, and has not logged in for ten days?
871. How long before a customer trusts the agent to send something externally without review?
872. Which roles need different onboarding paths, such as owner, manager, frontline staff, and contractor?
873. How do you prevent the agent from becoming one person's toy that never spreads to the rest of the team?
874. Where will customers go when they are stuck at 9 p.m.: docs, chat, a community, or the agent itself?
875. How will you measure time saved credibly, without inflating numbers the customer's CFO will challenge?
876. Which signals predict expansion, and which predict cancellation, in your first fifty accounts?
877. Can you describe the last customer who churned, why they left, and what would have kept them?
878. How will the agent handle a power user who wants advanced features before the rest of the team is ready?
879. Does the customer need to change any existing process to get value, or does the agent fit around it?
880. How many weeks until a customer reaches the outcome they bought, and what happens if they miss it?
881. Should onboarding include an interview where the agent learns the business, and how long can that take?
882. Who reviews the agent's first fifty outputs for a new account, your team or theirs?
883. How will you retrain users who learned bad habits, like pasting sensitive data into the wrong place?
884. Which unions, works councils, or HR policies could slow adoption in your larger target customers?
885. How will you communicate product changes so customers are not surprised when the agent behaves differently?
886. What incentives, if any, will you give a customer's staff for adopting the agent early?
887. If usage drops after the first month, which intervention will you try first and how will you measure it?
888. How do you handle a customer executive who mandates the agent while their team quietly works around it?
889. Which part of your current onboarding do customers skip, and is it actually necessary?
890. How will the agent ask for feedback without annoying users who are busy doing their jobs?
891. Will you run quarterly business reviews with customers, and what evidence will you bring to each one?
892. How do you win back a customer who left because the agent made an embarrassing mistake?
893. Which tasks should the agent hand back to a human permanently because adoption will never stick there?
894. How many support tickets per account per month can your margins absorb in the first year?
895. Should customers be able to pause the agent, and how will you make restarting it painless?
896. Who on your team is accountable for net revenue retention, and what number are they targeting?
897. How will the agent help a new employee at the customer's company get productive in their first week?
898. Which word do you use with staff to describe the agent: assistant, teammate, tool, or coworker, and why?
899. Have you watched a real customer onboard without helping them, and where did they get stuck?
900. When a customer finally says the agent is indispensable, what usually happened right before that?

<a id="team"></a>

## 19. Team, capital, and founder readiness

_An agent company is only as durable as the people and cash behind it. The builder asks these to test whether the founder has the bandwidth, the money, and the decision discipline to finish what they start._

901. Who writes the code today, and what happens to the roadmap if that person disappears for a month?
902. How many hours per week do you personally spend building, selling, and supporting, and which should you stop?
903. Which role do you need to hire next, and what would that person own on their first day?
904. Are you funding this with venture capital, customer revenue, consulting retainers, or savings, and why that mix?
905. How many months of runway do you have at current burn, counting only cash in the bank?
906. Which decisions do you make alone, which need a cofounder, and which need a board or advisor?
907. How do you decide between a feature a customer is begging for and the architecture work you know is overdue?
908. Who is your technical cofounder or principal engineer, and if none, how do you judge code quality?
909. What would make you shut this company down, and have you written that condition anywhere?
910. How much of the build will rely on AI coding agents, and who reviews their work before it reaches production?
911. Which consulting revenue keeps the lights on, and how do you stop it from consuming the product team?
912. Can you raise capital in the next twelve months, and what milestone would make investors say yes?
913. How will you compensate early hires: salary, equity, revenue share, or a mix, and what is your budget?
914. Which skills are missing from the founding team, such as sales, security, design, or operations?
915. How do you handle disagreements with your cofounder or key partner when the stakes are high?
916. How much personal financial risk can your household absorb, and does your family agree with that number?
917. Which tasks could a contractor do this month that you are still doing yourself?
918. How many separate businesses or projects are you running, and which one gets your best hours?
919. Who tells you the truth about your product when everyone else is being polite?
920. Should you hire a salesperson before you have a repeatable sales process, and what is your evidence?
921. How will you document decisions so a new hire understands why the system works the way it does?
922. What is your target gross margin after model costs, hosting, and support, and are you there today?
923. Which advisor or mentor has actually built and scaled an agent company, and how often do you talk?
924. How fast can you ship a fix to production, and who approves the release?
925. Which part of the business lives entirely in your head, with nothing written down?
926. How will you protect your focus when a big consulting client wants more of your time?
927. Does your cap table have room for a meaningful seed round and an employee option pool?
928. How do you decide when to buy a tool, use open source, or build it yourself?
929. Which metric would convince you to hire the next engineer, and which would convince you to wait?
930. When was the last time you said no to a good opportunity because it did not fit the plan?
931. How healthy are your sleep, exercise, and family time, and how do they show up in your decisions?
932. Who covers incidents at 2 a.m. when the agent sends something wrong to a customer?
933. How will you structure the company legally, and does it support raising money and issuing equity?
934. Which tasks in your own week could the agent you are building take over right now?
935. Are you willing to hire people smarter than you in their domain and let them overrule you?
936. How much budget is reserved for model API costs, and what happens if usage triples next month?
937. Which investor type fits this business: venture, angel, revenue-based financing, or none at all?
938. How will you onboard an engineer into a codebase with hundreds of migrations and decisions?
939. Who owns security and compliance in the company, even if it is a part-time responsibility?
940. Have you priced what it would cost to hire your current workload out to three people?
941. How do you keep shipping when you are also the salesperson, the support desk, and the accountant?
942. Which past founder mistake are you most at risk of repeating with this company?
943. How will you know the team is ready to take on a customer twice the size of your largest today?
944. Should you join an accelerator, and what specifically would you want from it beyond money?
945. What is the minimum team needed to support one hundred paying customers without breaking?
946. How often do you review the plan against actual results, and who holds you accountable for misses?
947. Which contractor or agency relationship is creating more risk than value right now?
948. If you got $2 million tomorrow, how would you spend the first $500,000?
949. How will you avoid building a company that only works because you personally work eighty hours a week?
950. Who would run the company for three months if you were unable to?

<a id="long-game"></a>

## 20. Learning loops, scale, and the long game toward AGI

_The only defensible agent business is one that gets better with every task and survives every model release. The builder asks these to see whether the founder has a compounding engine and an honest view of what could kill it._

951. How does every completed task make the agent measurably better for the next customer, not just the current one?
952. Which memories should the agent keep forever, which should expire, and who decides the retention rules?
953. How will you prevent one customer's lessons from leaking into another customer's agent?
954. What is the evaluation suite that tells you a new model or prompt change made the agent better, not worse?
955. Which domain will the agent expand into after its first job, and what data does it need first?
956. How will the agent learn from corrections, and how fast does a correction change its future behavior?
957. Should the agent ever modify its own instructions or tools, and what review gate sits in front of that?
958. If foundation model prices fall ninety percent, what happens to your pricing and your moat?
959. If OpenAI, Anthropic, or Google ships your core feature for free, why do customers still pay you?
960. Which proprietary data asset will you own in five years that no competitor can buy or scrape?
961. How will you turn the agent into a platform where third parties build skills, and when is that premature?
962. Ten years from now, what share of a customer's operations does your agent run, and what stays human?
963. Which regulation, if passed tomorrow, would hurt your business most, and how would you adapt?
964. How do you plan to stay model agnostic as providers change terms, prices, and capabilities every quarter?
965. Which failure mode could end the company in a single week: a data breach, a rogue action, or a lawsuit?
966. How will you measure compounding value, such as accuracy, autonomy, or time saved per account per quarter?
967. At what level of verified reliability will you let the agent take actions with no human approval?
968. How will the agent coordinate with other companies' agents when it negotiates, schedules, or buys on a customer's behalf?
969. Which experiments run continuously in production, and how do you stop a bad experiment from reaching customers?
970. What does your system do when the agent discovers a better way to do a task than the one the customer specified?
971. Who audits the agent's long-term behavior for drift, bias, or slow degradation that no single run reveals?
972. How will your architecture change between one hundred customers and one hundred thousand?
973. Which parts of your stack will you rebuild in three years, and are you designing for that replacement now?
974. Will you train or fine-tune your own models, and what volume of data would justify the cost?
975. How do you capture the reasoning behind human decisions, not just the decisions themselves, for future learning?
976. Is there a version of this company that becomes the operating system for small businesses, and is that the goal?
977. How will you handle the customer who wants to export all their memory and leave for a competitor?
978. Which ethical lines will the agent never cross, even if a customer pays more and asks directly?
979. How will you know if your agent is making customers dependent in a way that harms them?
980. Which acquirer would want this company in five years, and would you sell?
981. How does the agent learn a brand-new domain, like accounting or logistics, without a human writing every rule?
982. Are your benchmarks public, and would you publish results that show where the agent still fails?
983. How will the agent explain its past decisions years later when an auditor or court asks why it acted?
984. Which network effect, if any, makes the product better for everyone as more businesses join?
985. Should the long-term vision be one general agent or many specialist agents working as a team?
986. How much of your roadmap depends on model capabilities that do not exist yet, and what is plan B?
987. What would an AGI-grade version of your agent do that today's version cannot, in concrete business terms?
988. How do you keep institutional knowledge accurate when the customer's business, staff, and strategy keep changing?
989. Which open problems in agent reliability are you betting the company on solving yourself?
990. How will you protect the company against a single provider outage or a sudden API policy change?
991. When the agent becomes better than your own team at some tasks, how will you reorganize the company?
992. Which research partnerships, universities, or labs could accelerate your learning loop?
993. How do you decide when a customer-specific learning should become a default for every customer?
994. Could the agent eventually start and run a small business on its own, and would you want it to?
995. How will insurance, liability, and indemnity work when an autonomous agent causes real financial loss?
996. Which metric would tell you, five years in, that the company has truly built something durable?
997. How will you avoid shipping capability faster than customers can understand and govern it?
998. If your largest customer asked the agent to do something legal but harmful to their employees, what happens?
999. Do you have a written thesis on where agent technology will be in 2030, and how often do you revise it?
1000. If this company fails, what is the most likely reason, and what are you doing about it this month?

## Frequently asked questions

### What is an AGI business agent?

An AGI business agent is an AI agent that can run real parts of a company across many kinds of work, not just one narrow task. It plans, uses your tools, remembers what your organization knows, asks for approval when the stakes call for it, and leaves a receipt for every action it takes.

### What should you answer first before building an AI agent for your business?

Start with the job and the money. Name the one workflow the agent will own first, who it serves, what that work costs you today, and how you will prove the agent did it. Sections 1 through 5 of this list cover that ground.

### How long does it take to answer all 1,000 questions?

Most founders do not answer all of them in one sitting. Work one section per session. A focused team can cover the first five sections in a day and the full list over two to three weeks, with the answers becoming the agent's build brief.

### Why do autonomy and trust get their own sections?

Because they decide whether the agent gets used. An agent that acts without clear approval gates loses trust after one bad send, and an agent that cannot show evidence of its work gets ignored. Sections 10 and 11 set those rules before the build starts.

### Can I download the questions as a file?

Yes. The full list is published as a Markdown file at https://www.alexdly.ai/blog/agi-business-agent-questions/questions.md so you can paste it into your own docs, an AI assistant, or a project tracker.

### How does AlexDLY use these questions?

AlexDLY uses them to scope Alex, its AI execution layer, for each business. The answers become organizational memory, approval rules, and the first agenda of delegated work, so the agent starts from your reality instead of a generic template.

---

AlexDLY is the AI execution layer for founder-led teams. Build your free War Room at https://www.alexdly.ai/get-started
