Data Handling Guidelines for AI
General Guidelines
All members of the university community must access and use university data in ways that safeguard the data and protect the institution.
Units and members of the university community must ensure:
- Compliance with regulatory requirements, as well as third-party and other contractual data obligations.
- Data is used for the purposes for which it is collected and any restrictions for its use are observed.
- Data is collected, stored, and disposed of in ways appropriate to the risk and impact of unintended disclosure.
The following table provides guidance on the risks of using various tools with different types of university and external data.
'Personal' licenses, as described below, are truly personal and should carry those connotations, and in general should not be used by faculty and staff with university data. In contrast, it may be possible to acquire individual licenses under institutional or enterprise licensing agreements, which often include additional protections negotiated for the safe handling of university data.
Inclusion of a product in the list below does not necessarily mean an institutional agreement is in place that would allow for its use.
Definitions:
- Personal license: For private individuals who obtain an account or license for their own use, including free tiers
- Individual license: A single person's seat within an organization, obtained under the organization's enterprise agreement. These are covered by the "Institutional / Individual" rows in the table below.
- Institutional license: For organizations or business entities. Where an agreement is in place, its terms protect university data, and individual seats obtained under that agreement carry the same protections.
Legend
= Currently available
= Not currently available
= Recommended - Generally safe to use with minimal risk
= Warning - Can be used with caution and specific guidance
= Not Recommended - Significant risks that make the tool unsuitable for the data classification
Data Classification Reference - Data Classifications
Artificial Intelligence Tools - KB Article
Data Sanitization - KB Article
| LicenseType | Product | Data Classifications and Recommendations | |||
|---|---|---|---|---|---|
| Public Data | Internal Data | Limited Data | Restricted Data | ||
USask |
|||||
| Institutional / Individual | USask Data Centre Hosted AI | ||||
| Locally Hosted AI - Research Segment | 6 | 6 | |||
| Locally Hosted AI - Managed Computer | 6 | 6 | |||
| Locally Hosted AI - Unmanaged Computer | 6 | 6 | 6 | ||
Microsoft |
|||||
| Institutional / Individual 8 | Microsoft 365 Copilot Chat - web (Standard seat) | 7 | |||
| Microsoft 365 Copilot (Premium seat) | 7 | ||||
| Azure AI Services | 4 | 4 | 4 5 | ||
| Copilot Studio | 4 | 4 | 4 5 | ||
| Personal | Bing Search | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 |
| Copilot for Home | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
| Microsoft 365 Premium | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
AWS |
|||||
| Institutional / Individual 8 | Amazon Bedrock | 4 | 4 5 | 4 5 | |
| Amazon SageMaker AI | 4 | 4 5 | 4 5 | ||
| Other AWS AI services | 4 | 4 5 6 | 4 5 6 | ||
| Personal | AI Services | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 |
OpenAI |
|||||
| Institutional / Individual | ChatGPT Edu | 4 | 4 | 4 | |
| OpenAI API | 5 | 5 6 | 5 6 | ||
| Personal | ChatGPT Free | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 |
| ChatGPT Plus | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
|
|
|||||
| Institutional / Individual 8 | Gemini for Education (Standard seat) | 7 | |||
| Google AI Pro for Education (Premium seat) | 7 | ||||
| NotebookLM (included with Workspace for Education) | 7 | ||||
| Personal | Gemini Free | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 |
| Google AI Pro / Ultra | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
| NotebookLM | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
| NotebookLM premium | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
Anthropic |
|||||
| Institutional / Individual 8 | Claude for Education (Standard seat) | 7 | |||
| Claude for Education (Premium seat) | 7 | ||||
| Personal | Claude (Free) | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 |
| Claude Pro | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
Perplexity |
|||||
| Institutional / Individual | Perplexity Enterprise Pro | 4 | 4 6 | 4 6 | |
| Perplexity Deep Research | 4 6 | 4 6 | 4 6 | ||
| Sonar APIs by Perplexity | 5 | 5 6 | 5 6 | ||
| Personal | Perplexity (Free) | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 |
| Perplexity Pro | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 | |
DeepSeek |
|||||
| Personal | DeepSeek - cloud | 1 2 3 | 1 2 3 | 1 2 3 | 1 2 3 |
Zoom |
|||||
| Institutional / Individual 8 | Zoom AI Companion | ||||
*Notes : The AI marketplace is changing rapidly; as a result, the classification levels are illustrative and subject to change. Always consult current service terms or reach out to IT Support for additional consultation and support.
- Do not use Personal license services to process, store, or share university or third-party data as they lack the contracts or service agreements that safeguard ownership and control of university data.
- Personal license services may use submitted content to improve their models — some major vendors now train on consumer data by default, with multi-year retention, unless the individual opts out, and ICT cannot verify an individual's opt-out. Use of such services with USask data can result in the loss of IP/patent rights or the leaking of content to other unauthorized users.
- Be careful when publishing the outputs from Personal license services. These services do not include legal indemnification and publication of their outputs may lead to IP claims against generated content. Institutional licenses, in contrast, now generally include copyright indemnification, conditional on meeting the agreement's requirements (for example, leaving the vendor's built-in guardrails and content filters enabled) — confirm the specific terms of the applicable agreement.
- When training new models with sensitive USask data, take steps to restrict access permissions to the model. AI models may leak their underlying training data, so treat access to the model as equivalent to access to the raw underlying data.
- Use of advanced AI tools (APIs, agents, etc.) should be reviewed with ICT early in the planning stages. AI tools can leak underlying training data and produce confident fabrications, so great care should be taken before integrating AI tools into public-facing services.
- When using AI services with Restricted and Limited data, it is recommended to use USask-hosted platforms or services covered by a USask institutional agreement (Microsoft, AWS, Google Gemini for Education, Anthropic Claude for Education) with guidance from the Security and Operations teams.
- USask's institutional education agreements (Microsoft 365 Copilot, Google Gemini for Education / Google AI Pro for Education, and Anthropic Claude for Education — Standard and Premium seats) provide contractual data protections, including no training on institutional inputs and negotiated retention and security terms, sufficient for Public, Internal, and Limited data. Restricted data always carries caveats: its use with these services requires case-by-case review with ICT Security and Operations, and additional obligations (research agreements, health-data legislation, export controls) may prohibit use regardless of the vendor agreement. Note also that vendor retention and deletion commitments remain subject to legal process — a court order can require data to be preserved despite contractual deletion terms. Institutional agreements reduce but do not eliminate this exposure, which is a further reason personal licenses must not be used with university data.
- Data residency is determined by the specific agreement and tenant configuration, not by the license tier — an education SKU's data protections do not by themselves imply Canadian data residency. USask's Google Workspace for Education tenant is configured for Canadian data residency, and USask's Microsoft and AWS environments are provisioned in Canadian data centre regions. Note, however, that vendor residency commitments generally apply to customer data at rest — some services are not pinned to a region: Microsoft's commitments cover stored content (including stored Copilot interactions) while AI processing may occur outside Canada, and AWS keeps stored content in the selected region while features such as cross-region inference can move prompts and outputs out of region during processing unless constrained. For Restricted data subject to residency obligations, confirm both the at-rest and processing locations for the specific service with ICT before use.
Data Classification Guidance
USask Data
Type of USask Data
Institutional data: Data that is created, collected and stored by all units and members of the university community, in support of academic and administrative activities. Administrative data about teaching, learning, research and scholarly activity, such as grades, attendance, research grants held and publications generated, is considered institutional data.
Research data: Data that is created by or derived from research, scholarly, and artistic activities.
Personal data: Data that contains personal information about an identifiable individual as defined in the Provincial Local Authority Freedom of Information and Protection of Privacy Act (LAFOIP). This data if compromised or used inappropriately would have implications to the privacy of an individual.
Public Data
- Data that is (or can be) generally available to all employees, the general public, and the media. Unintended disclosure of such information has no effect on an individual, a group or institutional operations, assets or reputation.
- Examples include:
- Course catalogs
- University event schedules
- General contact information for departments
- Publicly available university policies and guidelines
Internal Data
- Data that is available to those members of the university community or research project team with a clear need for access as part of their employment, academic, or research duties and responsibilities. Unintended disclosure of such information has minimal or no effect on an individual, a group or institutional operations, assets or reputation.
- Examples include:
- Aggregated or de-identified personal information
- Intellectual property - patent applications and drafts of research papers
- Raw and processed data from research with no human/animal ethics considerations
- Internal meeting minutes and operational plans
Limited Data
- Data of a sensitive or confidential nature which is intended for limited internal use. Unintended disclosure of such information has moderate effect on an individual, a group or institutional operations, assets or reputation. Access to data is generally limited to individuals in specific job functions.
- Examples include:
- Student numbers or grades
- Detailed floor plans showing gas, water, hazardous materials
- Data on identifiable human biological material (e.g. tissue samples)
- Most identifiable personal information
Restricted Data
- Data of a highly sensitive or confidential nature which is intended for restricted internal use. Unintended disclosure of such information is serious and has severe or adverse effect on an individual, a group or institutional operations, assets or reputation. Access to data is restricted to specific legitimate use cases.
- Examples include:
- Confidential research data
- Data governed by third-party agreements/contracts with stipulations for restricted access
- Identifiable health information (must comply with relevant health information privacy laws)
- Data subject to export control regulations
- Social insurance numbers / Credit card numbers / Equity and Diverity Information such as Gender Identity
Third-Party Data
Is data that is created or owned by a third party and is being used in support of academic, research and administrative activities. This data if compromised or used inappropriately would have implications for the third party. This includes data such as licensed content, data sets, or copyrighted material.
Public (free)
- Examples may include:
- Web sites
- Open-access research articles
- Publicly available datasets
- Note: "Free" or "public" doesn't mean the content is not protected by copyright and it is not settled whether using copyrighted material in an AI tool is considered "fair use", verify compliance with organizational data sharing and copyright policies before using.
Limited / Restricted
- Data governed by third-party agreements/contracts/licenses. Examples may include:
- Confidential research data
- Identifiable health information
- Licenses data sets subject to intellectual property rights
- Data provided under a licensing agreement from the publisher (for example, published research papers available through the University Library)
- Data shared under non-disclosure agreements (NDAs - define scope of usage and sharing)