AI Governance

What MIT’s Generative AI Guidance Teaches Organisations About Data Risk

MIT’s guidance for generative AI offers a practical lesson for every organisation: classify data before choosing tools, and keep high-risk information out of them.

By Atul Singh8 min readSeptember 3, 2026
What MIT’s Generative AI Guidance Teaches Organisations About Data Risk

A generative AI policy should not begin with a list of approved prompts. It should begin with a question: what information are people allowed to submit, and to which tools?

MIT’s Guidance for use of Generative AI tools treats data classification as the foundation for responsible use. The guidance distinguishes between low-, medium- and high-risk data, requires MIT-licensed tools for institutional work involving low- and medium-risk data, and prohibits high-risk data from being entered into any generative AI tool — including licensed tools.

That is a useful operating principle beyond a university environment. For an SMB, the equivalent is simple: do not decide whether a tool is “safe” in the abstract. Decide whether it is appropriate for a particular type of business information.

The policy logic: tool choice follows data risk

Many organisations make the same mistake when adopting AI. They choose a popular tool, write a short acceptable-use policy, and expect employees to judge sensitive information correctly in the moment. That reverses the order of decisions.

MIT’s guidance uses data risk to determine what is permissible. In practical terms, the model looks like this:

Data category Typical examples Appropriate operating rule
Low risk / public Published webpages, public product descriptions, public job adverts May be used for low-risk tasks, subject to organisational policy
Medium risk / internal Internal process notes, non-public drafts, routine operational material Use an approved or licensed tool with suitable privacy agreements
High risk / sensitive Personal records, protected research, confidential client information, valuable intellectual property Do not enter into generative AI tools

The exact definitions must come from the organisation’s own legal, security and information-governance framework. MIT directs users to its Infoprotect resources to determine the relevant risk level and also highlights obligations under laws and policies such as FERPA and HIPAA, along with specific considerations for patents and research.

The important lesson is not that a licensed tool makes every use safe. MIT explicitly prohibits high-risk data even in its licensed tools. Licensing can reduce certain risks — such as third-party model training or external sharing where contractual protections apply — but it does not remove the need to classify the data first.

Why public tools create a separate risk

A publicly available AI service may use submitted content for model training or external sharing, depending on its terms and configuration. An employee who pastes an unannounced product specification, a client complaint containing personal information, or a draft legal agreement may therefore disclose more than intended.

The workflow problem is usually not malicious behaviour. It is ambiguity. An employee sees a task that would take ten minutes with AI and does not know whether the document is public, internal or sensitive. A workable policy must remove that ambiguity before the task reaches the prompt box.

A practical three-tier workflow for SMBs

An SMB can adapt the principle into a lightweight approval system without reproducing an institution’s entire governance structure. The following workflow is an implementation model derived from the risk-based approach in MIT’s guidance, not a substitute for legal or security advice.

1. Build a short data inventory

List the information employees regularly handle, rather than trying to classify every file in the business at once. Start with recurring workflows:

  • Sales proposals and discovery notes
  • Customer support tickets
  • Marketing content and campaign plans
  • Supplier and contractor records
  • HR documents and employee information
  • Product roadmaps and technical specifications
  • Financial forecasts and pricing models
  • Client contracts and legal correspondence

For each category, record who owns it, who can access it, and whether it includes personal, regulated, contractual or commercially confidential information.

2. Assign a default risk level

Use conservative defaults. Public marketing copy may be low risk. An internal campaign plan may be medium risk. A customer export containing names, contact details and purchase history should be treated as high risk until someone with appropriate responsibility determines otherwise.

The goal is not perfect classification. The goal is to make the safe decision the easy decision. If employees have to interpret a complicated matrix for every prompt, they will either avoid useful tools or ignore the policy.

3. Map each tier to an approved action

Create a one-page internal rule set such as:

  • Public data: approved public tools may be used for drafting, summarising or transforming the material.
  • Internal data: use only the organisation’s approved, contractually governed tools; do not upload information from other clients or departments without permission.
  • Sensitive or high-risk data: do not submit it to a generative AI tool. Use ordinary internal systems or an approved specialist workflow instead.

This turns policy into a decision employees can apply during real work.

4. Separate the task from the source document

Often the employee does not need to upload the original material. They may only need help with a general transformation. For example:

  • Instead of uploading a customer complaint, remove names, order numbers and identifying details, then ask for a neutral response structure.
  • Instead of pasting a confidential proposal, describe the industry, buyer role and desired tone using invented or sanitised examples.
  • Instead of submitting a full employee policy, ask for a checklist based on a public template and review the result internally.

Anonymisation is not automatically sufficient. Small datasets, unusual facts or combinations of details can still identify a person or client. Treat sanitisation as a risk-reduction step, not permission to ignore the classification rules.

5. Put tool approval through procurement

MIT’s guidance notes that acquiring new AI tools should follow the relevant procurement process so vendor agreements can be assessed. The same principle applies to a small business, even if the process is a single owner and a short review checklist.

Before approving a tool, check:

  • Whether submitted data is used for model training
  • Whether the vendor shares data with third parties
  • Where data is stored and processed
  • How deletion requests work
  • Whether administrators can control retention and access
  • Whether the contract addresses confidentiality and regulated information
  • What happens to data when the subscription ends

Do not rely only on a product webpage or a checkbox in the sign-up flow. Keep the vendor’s terms, internal decision and approved use cases together so the policy can be reviewed later.

Applying the model to common business decisions

Consider a small recruitment firm preparing a candidate shortlist. A public AI tool may help rewrite a job advert based on information already published. It should not receive candidate CVs, interview notes or health-related information unless the organisation has a specifically approved and legally compliant workflow. Under a strict high-risk rule, those materials remain outside generative AI tools altogether.

Now consider a software company preparing release notes. Publicly documented features can be drafted in an approved tool. Internal bug summaries may require a licensed tool with suitable contractual protections. An unreleased security vulnerability, customer-specific configuration or patent-related technical detail should be treated as sensitive and kept out of general-purpose AI tools.

For a professional services firm, the distinction may arise within one document. A client report can contain public market context alongside confidential recommendations and personal information. “The report is internal” is not a sufficient classification. The firm should separate the public background from the sensitive client material before considering AI assistance.

These examples show why a blanket rule such as “AI is allowed for low-risk work” needs operational definitions. Employees need to know what low risk looks like in their own workflows, which tools are approved, and when the answer is simply no.

Callout: Make the default rule visible

Put the three data tiers and their permitted tools next to the systems where work happens — such as the document platform, CRM or internal knowledge base. A policy buried in an employee handbook will not guide a decision made in thirty seconds.

Failure modes to address before rollout

A risk-based policy is useful, but it can still fail if implementation is careless.

Treating a vendor contract as a universal safety guarantee. A privacy agreement may limit model training or external sharing, but it does not make high-risk data appropriate for the tool. It also does not solve poor access controls, accidental disclosure or incorrect outputs.

Classifying tools instead of data. Saying “we use the secure AI platform” encourages employees to submit everything. The correct question remains: what is this information, and what rules apply to it?

Allowing sensitive data after informal redaction. Removing a name while leaving a unique role, date, location and case details may still expose an identifiable person or client.

Ignoring legal and contractual obligations. Privacy law, sector regulation, research rules, intellectual property obligations and client contracts may impose stricter requirements than an internal AI policy.

Failing to plan for mistakes. MIT’s guidance directs users who have entered high-risk data into a generative AI tool to contact security immediately. An SMB should define its own equivalent escalation route: who must be told, what evidence to preserve, how access is revoked and when affected customers or regulators must be considered.

Writing a policy no one can apply. A long document that does not identify approved tools, prohibited data and escalation contacts is not operational guidance. Start with a one-page rule set, then expand it as real cases reveal gaps.

A sensible starting point

An organisation does not need to automate every workflow before it can use generative AI responsibly. Start with three actions:

  1. Classify the five to ten information types employees use most often.
  2. Approve specific tools for low- and medium-risk work after reviewing their data practices.
  3. Make high-risk data a clear prohibition, with an incident route for accidental submission.

Review the classifications quarterly and whenever a new tool, regulation, client contract or business process is introduced. Also verify current tool availability and definitions rather than assuming that a platform or policy remains unchanged.

The deeper lesson from MIT’s guidance is straightforward: responsible AI adoption is an information-governance problem before it is a prompt-writing problem. Organisations that decide what data may enter which tool can expand useful adoption without asking employees to make security judgments from vague assurances.

FAQs

Does using a licensed or enterprise AI tool make sensitive data safe to submit?

No. MIT’s guidance prohibits high-risk data in any generative AI tool, including licensed tools. Licensing may provide stronger privacy protections, but organisations must still classify data and follow legal, contractual and institutional requirements.

What should an organisation do if someone submits high-risk data to an AI tool?

Create an immediate escalation process. Record what was submitted, when and to which tool; notify the designated security or privacy contact; preserve relevant evidence; and follow applicable contractual, legal and incident-response requirements.

How can a small business begin classifying data for AI use?

Start with the information used in the most common workflows, such as customer records, proposals, HR files and product plans. Assign public, internal or sensitive defaults, then map each category to approved tools and prohibited actions.

A

Atul Singh

15 years across teaching, sales, and building. Trained 2,500+ students. Six years in corporate sales and social media. Six years building web and AI products for SMBs at Qriyas. Based in Noida, working with sales and marketing professionals across the US, UK, Australia, and English-speaking markets globally.