📅 Data checked: Aug 28, 2026 · Today: · Terms can change at any time. Always check a database’s current terms before you rely on them.
NCCUDatabases, copyright & AI
NCCU Library · a database & AI guide for students and staff

Download, mine, feed to AI:
three different rights.

(“Download” means saving a file. “Mining” (TDM) means letting a computer read and analyze many texts at once. “Feed to AI” means giving content to a tool like ChatGPT.)

You can still use AI legally. This page shows you how to do it safely. It answers one question for NCCU students and staff: can you give an article from a library database to an AI tool to summarize, translate, or rewrite? Below, you can check the terms, run the self-check, and read the legal basis.

The one-minute answer
  • OKUse a database’s own built-in AI, such as Scopus AI or JSTOR. (The Library has a separate page on these.) You can also use your own writing, public-domain works, or open-access content.
  • CarefulPasting a full-text subscribed article into an outside AI like ChatGPT usually breaks the database’s terms. Check the terms first, or paste only key points or your own notes.
  • Don’tBulk-download full text with crawlers or bots. Nearly all 19 databases forbid this. Also, don’t give AI any file that is pirated or from an unknown source.
  • Not sure?Use the self-check below. It gives you an answer in a few taps.
Self-check

Self-check: can I upload this to AI?

Start with a quick self-check. Pick your situation, and you’ll get advice right away. The reasons come further down.

SELF-CHECK

Safer ways to do it

In practice: from what’s public so far, people are rarely sued just for pasting a single PDF into AI to study it, and personal noncommercial research has some room to claim fair use. But “rarely sued” is not the same as “clearly legal.” And breaking a contract is a separate question from whether you can claim fair use under copyright.

1

Use a licensed tool that won’t train on or keep your data

Enterprise or education editions that clearly don’t store or train on your content carry the least risk. Avoid free tools that keep or harvest what you enter.NSTC guidelines ↗

2

Use the database’s built-in AI features

Such as the ScienceDirect Reading Assistant, Ovid, or the ProQuest Research Assistant. The publishers have already licensed these.NCCU Library · database AI features ↗

3

Use open-access or CC-licensed content

Still check the CC terms. NC or ND licenses may not work with a commercial AI platform.Creative Commons ↗

4

Don’t upload the full text

Paste your own notes, a few short quotes, or the key points — not the whole PDF.MODA handbook ↗

5

Get permission when you need it

If you’re not sure what a database allows, check with the publisher or platform, or request a license.Down to the terms ↓

6

Pause before you upload

Before you attach a file, check whether it involves copyright, privacy, or confidential material.MODA handbook ↗

When a database says no: what you can legally do

Often people need a lot of text for digital humanities, bibliometrics, or model training — not to pirate anything. Instead of scraping it yourself or dropping full text into an outside AI, here are legal paths, organized as “what you want to do → which route to take.” (What’s actually available still depends on your institution’s license and each platform’s terms.)

Want to bulk-download or write a crawler?

→ Use official APIs and cloud analysis platforms
  • Official TDM APIs: lawful subscribers can get an API key from the Elsevier ScienceDirect API and pull full text and abstracts within the rate limits; the Crossref TDM API gives unified metadata and TDM access across many journals — far lower risk than scraping yourself.
  • Cloud analysis (no need to download full text locally): if your school licenses ProQuest TDM Studio, use its interface or Jupyter notebooks for large-scale mining; JSTOR also offers an official text-analysis service (Constellate closed in mid-2025; now JSTOR Text Analysis Support). HathiTrust HTRC uses non-consumptive research: the algorithm is sent into a secure server to run, or only statistical “extracted features” are released, so copyrighted full text never leaves.

Training a model, doing RAG, or building a local knowledge base?

→ Use clearly open-licensed data
  • OpenAlex: 250M+ scholarly metadata records, open with a free API, and a community MCP already exists — currently the best option for letting AI search scholarly context legally.
  • PubMed Central OA Subset / arXiv / Europe PMC: millions of open full-text articles cleared for machine reading.
  • CORE / Unpaywall: to find a specific paywalled paper, use these to locate a lawful Green OA (author self-archived) version.

Just want AI to help synthesize or organize?

→ Use metadata + your own notes, not the whole PDF
  • Metadata instead of full text: export a bibliographic file from WOS/Scopus (RIS/CSV with titles, abstracts, keywords) and hand those abstracts — which raise no full-text reproduction issue — to AI for topic clustering or trend analysis.
  • Citation-note method: jot down the core ideas and methods in Zotero/EndNote, and give AI only your own notes or short excerpts, sidestepping reproduction and contract concerns.

Really must analyze full text with AI?

→ Use the publisher’s built-in AI, or a local offline model
  • Publishers’ built-in AI: like Scopus AI or the Web of Science Research Assistant — already licensed by the publisher, with citations you can trace; the easiest and most compliant option.
  • Local offline models (Ollama / LM Studio): run on your own computer offline, so nothing is sent to a third-party server (no third-party disclosure) — more defensible than uploading to OpenAI’s or Anthropic’s public cloud.
Researchers, take special note — never feed a manuscript under review or unpublished into AI: if you’re helping with peer review or serving as an editor, don’t upload an unpublished manuscript under review, its data, or its figures into any public or unlicensed generative AI. COPE, ICMJE, and the major publishers (Elsevier, Springer Nature, Wiley, SAGE, and others) expressly forbid this, for two reasons: ① review material is confidential, so uploading it to a public-cloud AI is a disclosure to a third party; ② if an author’s unpublished work is logged by an AI or folded into training, it can seriously harm their rights. Dropping a paper under review into AI to draft your review is treated as a serious research-ethics violation and a breach of confidentiality.

For any one database, three separate sets of rules usually apply: one for bulk downloading, one for text and data mining (TDM), and one for sending content to outside AI tools. This table sorts the public terms of 19 common academic and news databases into those three columns. You see the bottom line first, then the source text.

Bulk download & scraping
Almost always banned.
None of the 19 databases let you mass-download with crawlers or bots.
Text & data mining (TDM)
“Conditional” at best.
Even then, you need lawful access, a noncommercial purpose, and an approved API or contract.
External AI / RAG
None allow it outright.
At best, you need a separate license. Most simply forbid it.
i

This is a summary of terms, not legal advice. Your school’s signed license agreements, a product’s own terms, and open-access licenses (like CC BY) may add to or override what you see here. Check with the publisher or platform before you rely on it.

Comparison

What each database allows

Tap any row to open it. You’ll see plain-language notes, the source clauses, and official links. (“TDM” means letting a computer read and analyze many texts at once, instead of one by one.)

Five things to check before you read the table
  1. What kind of access do you have? A campus subscription, a personal account, open access, and an API each allow different things.
  2. Do you need the full text or just metadata? Being able to get DOIs or abstracts doesn’t mean you can get full-text PDFs.
  3. Is this noncommercial research? Most TDM exceptions apply only to noncommercial research.
  4. Will the data leave the approved place? Uploading it to an outside AI or a vector store makes a new copy and shares it.
  5. Does the contract mention AI or RAG? A standard TDM license does not automatically cover training, RAG, or fine-tuning.
Prohibited: clearly banned, with no general way to get a license Conditional: allowed if you meet the conditions (lawful access, noncommercial, an approved route) License required: not allowed by default, but you can ask for or contract one Unspecified: the public terms don’t clearly say