The risky moment in an email-agent purchase is not when the model writes a bad sentence. It is when a vendor receives access to your inbox, CRM, customer data and sending identity before your team has defined what the agent is allowed to do.

That access can be useful. A good system can classify inbound mail, draft responses, route leads, prepare follow-ups, update records and surface exceptions. But the same integration can also create deliverability problems, expose sensitive data, send a message under the wrong identity, or make it difficult to reconstruct why an action happened.

So the buying question is bigger than “which model is smartest?”

Before a vendor or implementation partner touches production email, ask the questions below. The list is intentionally operational. A strong vendor should be able to answer with architecture, controls, logs and examples—not only a demo.

1. What exactly can the agent read?

Start with data scope.

Ask the vendor to list every object the system can access:

  • message body;
  • attachments;
  • contact records;
  • calendar context;
  • CRM fields;
  • support history;
  • internal notes;
  • files linked from a message;
  • authentication tokens;
  • sending-domain configuration.

Then ask whether access is tenant-wide, mailbox-specific, label-specific or workflow-specific.

A useful answer sounds like: “The agent reads only messages in this shared inbox and only these CRM fields.” A weak answer sounds like: “We connect to your workspace and use what is needed.”

Red flag: broad access is required even when the workflow is narrow.

2. Can permissions be separated by action?

Reading, drafting and sending should not automatically be one permission.

Look for separate controls for:

  • read;
  • classify;
  • create draft;
  • edit CRM;
  • schedule;
  • send;
  • delete;
  • archive;
  • change labels;
  • trigger another system.

For higher-risk workflows, the agent should be able to prepare an action without executing it.

Question to ask: can we require human approval for send, while still allowing automated classification and drafting?

3. Where is customer data stored, and for how long?

You need a data-flow diagram, not a sentence saying “enterprise-grade security.”

Ask:

  • where message content is processed;
  • where logs are stored;
  • retention period;
  • whether prompts or content are reused for model training;
  • how backups are handled;
  • whether deleted customer data is removed from derived stores;
  • what subprocessors are involved.

The correct answer depends on your organization, contracts and jurisdiction. The buying discipline is to make the path visible.

4. Which model or models are used, and can they change without notice?

Email-agent vendors often orchestrate multiple models.

That may be sensible, but it affects cost, behavior, data handling and reproducibility.

Ask whether the system uses:

  • one fixed model;
  • model routing;
  • customer-selected models;
  • fallback models;
  • local or third-party models;
  • vendor fine-tunes.

Then ask what happens when a model version changes.

A good contract should define whether material behavior changes are announced and whether you can test them before production.

5. How does the agent handle instructions inside an email?

An email is untrusted input.

A customer, spammer or attacker can write text that looks like an instruction to the AI. If the agent treats message content as trusted operational policy, it can be manipulated.

Ask the vendor how it separates:

  • customer content;
  • system instructions;
  • internal policies;
  • tool permissions;
  • retrieved knowledge.

NIST’s Generative AI Profile is useful here because it treats generative-AI risk as a management problem across design, deployment and monitoring—not simply a model-accuracy question.

Red flag: “The model can tell which instructions are malicious” is the entire control story.

6. Can we see why an action happened?

You need an audit trail.

For every important action, the system should be able to record:

  • input event;
  • workflow version;
  • model or rules used;
  • retrieved context;
  • proposed action;
  • approval status;
  • final sender identity;
  • external tool call;
  • result or error;
  • timestamp.

The goal is not to expose hidden model reasoning. It is to reconstruct the operational chain.

If a customer complains about an email, somebody should be able to answer: who or what sent it, under which rule, using which data?

7. How is sending identity protected?

Ask whether the agent can send from:

  • personal mailboxes;
  • shared inboxes;
  • aliases;
  • transactional systems;
  • marketing platforms.

Then ask how the system prevents the wrong identity from being used.

The Gmail and Yahoo sender requirements make the infrastructure side important too. Authentication, complaint rates and unsubscribe behavior can affect delivery. An email-agent project cannot treat deliverability as a cosmetic afterthought.

8. Who owns SPF, DKIM and DMARC configuration?

Do not accept “your IT team handles that” without a responsibility map.

The agent vendor may not control DNS, but it should clearly state what it expects.

Google’s Gmail guidance says bulk senders to personal Gmail accounts must meet authentication and unsubscribe requirements, and Google’s current FAQ explains how bulk-sender status and spam-rate enforcement work. Yahoo likewise requires stronger authentication and easy unsubscribe for bulk senders and emphasizes low complaint rates.

Ask who verifies that the configuration is actually working after setup.

9. How are unsubscribe and suppression requests enforced?

A robust system should have a suppression layer that the agent cannot casually override.

For promotional mail, ask:

  • where unsubscribe status lives;
  • how quickly it propagates;
  • whether all sending paths respect it;
  • how one-click unsubscribe is handled where required;
  • whether a human can accidentally resend to a suppressed recipient.

The FTC’s CAN-SPAM guidance also matters for U.S. commercial email. It requires accurate header information, non-deceptive subject lines, an opt-out mechanism and honoring opt-out requests.

Do not turn this into “CAN-SPAM means cold email is always allowed.” Laws differ by jurisdiction and context. The vendor should support your compliance controls, not provide blanket legal permission.

10. What complaint-rate threshold stops automation?

Google recommends keeping spam rates low and specifically warns bulk senders to keep user-reported spam below 0.1% and prevent it from reaching 0.3% or higher. Yahoo’s sender guidance also uses 0.3% as a key threshold.

Ask the vendor:

  • can complaint data be ingested;
  • can campaigns or workflows pause automatically;
  • who receives the alert;
  • what is the recovery procedure?

A system that can send thousands of messages should be able to stop itself when reputation deteriorates.

11. What happens when the agent is uncertain?

The best answer is not “it never hallucinates.”

Look for a confidence or exception policy.

Examples:

  • missing order number → ask for clarification;
  • legal threat → escalate;
  • refund above limit → human approval;
  • conflicting CRM data → do not send;
  • attachment type unsupported → route to manual review.

A useful email agent knows when not to act.

12. Can the vendor prove the approval workflow cannot be bypassed?

A visual “Approve” button is not enough.

Ask whether enforcement happens in the backend. Can a workflow update, API call or alternate sending path bypass the approval requirement?

You want authorization controls, not merely interface conventions.

13. How are templates, policies and knowledge sources versioned?

If your sales policy changes on Monday, you need to know which messages were generated using the old version.

Ask for version history on:

  • prompts;
  • workflow rules;
  • approved templates;
  • knowledge bases;
  • pricing tables;
  • escalation policies.

Operational systems need change control.

14. What is the real unit cost?

Do not price only the model tokens.

Include:

  • vendor platform fee;
  • mailbox or seat fees;
  • model usage;
  • enrichment;
  • CRM automation;
  • deliverability tooling;
  • implementation;
  • approval labor;
  • monitoring;
  • exception handling;
  • failed-message investigation.

A cheap model can sit inside an expensive operating system.

15. What breaks when our volume doubles?

Ask about limits before you hit them.

What happens to:

  • API quotas;
  • mailbox limits;
  • queue latency;
  • model rate limits;
  • CRM write capacity;
  • log retention;
  • approval workload?

The answer should include backpressure and graceful failure, not only an upgraded pricing tier.

16. Can we export our data, prompts, logs and suppression lists?

Exit is part of procurement.

Ask what you can export if you leave:

  • contacts;
  • message history;
  • workflow definitions;
  • templates;
  • prompt versions;
  • audit logs;
  • suppression data;
  • performance metrics.

If critical state exists only inside the vendor’s interface, switching cost can become operational risk.

17. Who is responsible during an incident?

Ask for the escalation path before launch.

You need named responsibilities for:

  • accidental mass send;
  • compromised credential;
  • data exposure;
  • deliverability collapse;
  • wrong-domain sending;
  • model behavior incident;
  • vendor outage.

“How do we contact support?” is not the same as an incident plan.

18. What will success look like after 30 days?

The vendor should help you define operational metrics, not just promise “more productivity.”

A useful pilot scorecard can include:

  • draft acceptance rate;
  • human edit rate;
  • time to first response;
  • escalation rate;
  • send-error rate;
  • complaint rate;
  • unsubscribe rate;
  • qualified-reply rate;
  • cost per handled conversation;
  • incidents requiring manual recovery.

The goal is to decide whether the agent is reducing work without creating hidden risk.

A simple procurement scorecard

Score each vendor from 0 to 2:

  • 0 — vague or unavailable
  • 1 — available with manual process or limitation
  • 2 — enforced, observable and exportable

Use the score across six categories:

Category Weight
Data and permissions 25%
Deliverability and compliance controls 20%
Approval and exception handling 20%
Auditability and change control 15%
Economics and scale 10%
Portability and incident response 10%

Do not let a polished demo compensate for a zero in permissions, suppression or auditability.

An email agent is not only a writing tool. Once it can read customer data and act through your sending identity, it becomes part of your operating infrastructure.

Buy it with the same seriousness.

Sources

Related Reading