Download, mine, feed to AI:
three different rights.
(“Download” means saving a file. “Mining” (TDM) means letting a computer read and analyze many texts at once. “Feed to AI” means giving content to a tool like ChatGPT.)
You can still use AI legally. This page shows you how to do it safely. It answers one question for NCCU students and staff: can you give an article from a library database to an AI tool to summarize, translate, or rewrite? Below, you can check the terms, run the self-check, and read the legal basis.
- OKUse a database’s own built-in AI, such as Scopus AI or JSTOR. (The Library has a separate page on these.) You can also use your own writing, public-domain works, or open-access content.
- CarefulPasting a full-text subscribed article into an outside AI like ChatGPT usually breaks the database’s terms. Check the terms first, or paste only key points or your own notes.
- Don’tBulk-download full text with crawlers or bots. Nearly all 19 databases forbid this. Also, don’t give AI any file that is pirated or from an unknown source.
- Not sure?Use the self-check below. It gives you an answer in a few taps.
Self-check: can I upload this to AI?
Start with a quick self-check. Pick your situation, and you’ll get advice right away. The reasons come further down.
Safer ways to do it
In practice: from what’s public so far, people are rarely sued just for pasting a single PDF into AI to study it, and personal noncommercial research has some room to claim fair use. But “rarely sued” is not the same as “clearly legal.” And breaking a contract is a separate question from whether you can claim fair use under copyright.
Use a licensed tool that won’t train on or keep your data
Enterprise or education editions that clearly don’t store or train on your content carry the least risk. Avoid free tools that keep or harvest what you enter.NSTC guidelines ↗
Use the database’s built-in AI features
Such as the ScienceDirect Reading Assistant, Ovid, or the ProQuest Research Assistant. The publishers have already licensed these.NCCU Library · database AI features ↗
Use open-access or CC-licensed content
Still check the CC terms. NC or ND licenses may not work with a commercial AI platform.Creative Commons ↗
Don’t upload the full text
Paste your own notes, a few short quotes, or the key points — not the whole PDF.MODA handbook ↗
Get permission when you need it
If you’re not sure what a database allows, check with the publisher or platform, or request a license.Down to the terms ↓
Pause before you upload
Before you attach a file, check whether it involves copyright, privacy, or confidential material.MODA handbook ↗
When a database says no: what you can legally do
Often people need a lot of text for digital humanities, bibliometrics, or model training — not to pirate anything. Instead of scraping it yourself or dropping full text into an outside AI, here are legal paths, organized as “what you want to do → which route to take.” (What’s actually available still depends on your institution’s license and each platform’s terms.)
Want to bulk-download or write a crawler?
- Official TDM APIs: lawful subscribers can get an API key from the Elsevier ScienceDirect API and pull full text and abstracts within the rate limits; the Crossref TDM API gives unified metadata and TDM access across many journals — far lower risk than scraping yourself.
- Cloud analysis (no need to download full text locally): if your school licenses ProQuest TDM Studio, use its interface or Jupyter notebooks for large-scale mining; JSTOR also offers an official text-analysis service (Constellate closed in mid-2025; now JSTOR Text Analysis Support). HathiTrust HTRC uses non-consumptive research: the algorithm is sent into a secure server to run, or only statistical “extracted features” are released, so copyrighted full text never leaves.
Training a model, doing RAG, or building a local knowledge base?
- OpenAlex: 250M+ scholarly metadata records, open with a free API, and a community MCP already exists — currently the best option for letting AI search scholarly context legally.
- PubMed Central OA Subset / arXiv / Europe PMC: millions of open full-text articles cleared for machine reading.
- CORE / Unpaywall: to find a specific paywalled paper, use these to locate a lawful Green OA (author self-archived) version.
Just want AI to help synthesize or organize?
- Metadata instead of full text: export a bibliographic file from WOS/Scopus (RIS/CSV with titles, abstracts, keywords) and hand those abstracts — which raise no full-text reproduction issue — to AI for topic clustering or trend analysis.
- Citation-note method: jot down the core ideas and methods in Zotero/EndNote, and give AI only your own notes or short excerpts, sidestepping reproduction and contract concerns.
Really must analyze full text with AI?
- Publishers’ built-in AI: like Scopus AI or the Web of Science Research Assistant — already licensed by the publisher, with citations you can trace; the easiest and most compliant option.
- Local offline models (Ollama / LM Studio): run on your own computer offline, so nothing is sent to a third-party server (no third-party disclosure) — more defensible than uploading to OpenAI’s or Anthropic’s public cloud.
For any one database, three separate sets of rules usually apply: one for bulk downloading, one for text and data mining (TDM), and one for sending content to outside AI tools. This table sorts the public terms of 19 common academic and news databases into those three columns. You see the bottom line first, then the source text.
None of the 19 databases let you mass-download with crawlers or bots.
Even then, you need lawful access, a noncommercial purpose, and an approved API or contract.
At best, you need a separate license. Most simply forbid it.
This is a summary of terms, not legal advice. Your school’s signed license agreements, a product’s own terms, and open-access licenses (like CC BY) may add to or override what you see here. Check with the publisher or platform before you rely on it.
What each database allows
Tap any row to open it. You’ll see plain-language notes, the source clauses, and official links. (“TDM” means letting a computer read and analyze many texts at once, instead of one by one.)
- What kind of access do you have? A campus subscription, a personal account, open access, and an API each allow different things.
- Do you need the full text or just metadata? Being able to get DOIs or abstracts doesn’t mean you can get full-text PDFs.
- Is this noncommercial research? Most TDM exceptions apply only to noncommercial research.
- Will the data leave the approved place? Uploading it to an outside AI or a vector store makes a new copy and shares it.
- Does the contract mention AI or RAG? A standard TDM license does not automatically cover training, RAG, or fine-tuning.
Understanding the rules: copyright, court cases, and how countries handle it
Terms often say “follow copyright law,” but copyright is a separate set of rules. This section explains the two layers — the contract and copyright — then walks you through 2026 court cases and the rules in different countries. What follows is general information to help you judge the direction — not legal advice for your specific case; when in doubt, consult a professional.
The database license terms
This is usually the most direct limit. As the table above shows, most database terms forbid or restrict handing full text to an outside AI. Even if fair use might apply under copyright, the contract can still restrict it. The library buys a license — you don’t personally own the content. But the contract layer can be pushed back on, too: in the U.K., Jisc (which negotiates for universities) urged them in 2024 to refuse overly restrictive AI clauses and protect staff and students’ lawful research and noncommercial mining.
Reproduction, adaptation, fair use
When you upload full text to AI, you usually make a copy and hand it to a third party — before the AI even starts to summarize. Under Taiwan’s Copyright Act, this is usually an act of reproduction (Art. 3). Whether it counts as infringement depends on whether you have permission or can claim fair use (Arts. 44–65, decided case by case). It’s even clearer internationally. Japan is often called the most permissive place for AI, because it has an “information analysis” exception (§30-4). But in 2024 Japan’s Agency for Cultural Affairs made clear: if a database was made to be sold for analysis, then scraping it without permission — or vectorizing it for RAG — still counts as “unreasonably harming the rights holder’s interests,” so it is not exempt. In other words, a database license holds up even against the most permissive copyright exception.
Three uses, from lower to higher risk
A rule of thumb: if what the AI produces looks a lot like the original, it infringes. A summary of plain facts or ideas is usually fine. The closer you get to a word-for-word translation or a copied rewrite, the riskier it gets.
Reading help and summaries
The main risk is that uploading usually involves making a copy. A high-level summary you make just to understand something yourself carries less risk on the output side, because copyright protects the expression, not the ideas. But uploading the full text is still a copy and a disclosure to a third party.
Art. 3 reproduction · Art. 10-1 idea/expressionTranslation
Translation is generally treated as an “adaptation,” a right that belongs to the author, and the result is generally a derivative work. You have more room for personal use. But once you share it — or the translation closely reproduces the original’s wording — you need permission.
Art. 28 adaptation · Art. 3(1)(11)Rewriting
Rewriting is usually an adaptation, or derivative work. The closer it comes to replacing the original, the worse it looks under the fourth fair-use factor (effect on the market) — especially if you share or reuse the result.
Art. 28 adaptation · Art. 65 market effectKey cases (including Taiwan’s first): how you got the content matters most; even noncommercial, open-source sharing can infringe if you reproduce content in bulk without permission (open for the cases)
First, to be clear: the foreign cases mostly sue AI companies over how they train their models, and don’t bind Taiwan’s courts — they just show the international debate. But Taiwan now has its first case — against an individual, and criminal — and that’s the one to read.
CNA v. an NTU PhD student (Taiwan’s first)
To fill a gap in Traditional Chinese training data, an NTU PhD student compiled web text into an open dataset, “fineweb-zhtw,” and shared it free on Hugging Face — but it contained about 140,000 unauthorized news items from Taiwan’s Central News Agency (CNA) (2011–2021). In July 2025 CNA filed a criminal copyright complaint; the two sides later settled, so no court ruled.
Bartz v. Anthropic ($1.5 billion settlement)
The largest copyright settlement in U.S. history. In a partial summary-judgment ruling, Judge Alsup initially suggested that training AI on lawfully obtained books may have room for fair use, while downloading them from a pirate library infringes. But the case then settled, so this view was never reviewed on appeal and is not yet nationwide precedent; the settlement itself covers only how the content was obtained (the input side).
Elsevier et al. v. Meta
Five major academic and educational publishers (Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill), together with author Scott Turow, are suing Meta and Mark Zuckerberg. They say Meta downloaded millions of books and journal articles from pirate (shadow) libraries — and scraped others from behind paywalls — to train Llama, then removed copyright management information (CMI) to hide the source. This is the first copyright suit publishers have brought against an AI company (case no. 1:26-cv-03689).
New York Times v. OpenAI / Microsoft
The New York Times is suing OpenAI and Microsoft, claiming they used millions of its articles without permission to train ChatGPT, and that ChatGPT can reproduce its reporting almost word for word, hurting subscriptions. In April 2025 the court let the core copyright claims proceed; in January 2026 it ordered OpenAI to hand over 20 million ChatGPT conversation logs. Case No. 1:23-cv-11195, still in litigation.
Dow Jones / New York Post v. Perplexity
Dow Jones (parent of the Wall Street Journal) and the New York Post are suing the AI answer engine Perplexity, claiming it scrapes past the paywall, stores copyrighted articles in a vector database for RAG, reproduces much of the original in its answers, and lets users “skip the links” — never clicking through to the source — creating market substitution. Case No. 1:24-cv-07984, in litigation.
What Taiwan’s official guidance says
The Ministry of Digital Affairs (MODA) Reference Handbook on AI Use in the Public Sector (Jan 2026) and the National Science and Technology Council (NSTC) Guidelines on Using Generative AI (Oct 2023) are written for government agencies. But the same principles apply to students and staff, and they back up the advice on this page.
- Which laws apply: the government lists the Artificial Intelligence Basic Act, the Personal Data Protection Act, the Copyright Act, the Trademark Act, the Patent Act, and the Trade Secrets Act. (Handbook 1.6)
- Keep the data clean: data you train on or use should leave out anything that infringes copyright or personal data. (Handbook checklist 2.3.2)
- Off-the-shelf AI “eats” your data — don’t feed it secrets: data you give a commercial AI service “may be reused directly by the provider.” Don’t enter personal or confidential data into a public or API-connected service. Write confidential documents yourself and don’t use generative AI for them. (Handbook 3.1.3, 4.3.6; NSTC guidelines)
- Follow the service terms: when you use an AI service, follow its terms of use to stay within the law. (Handbook 4.2.7)
- Don’t treat AI as an expert, and don’t fully trust it: in fields that demand precision, like law and medicine, don’t use generative AI as a substitute for an expert. A person must make the final professional call; don’t rely on the output blindly. (Handbook p.9; NSTC guidelines)
Taiwan’s other AI-related laws and guidance (from the MODA handbook, p.20)
The following is drawn from the Ministry of Digital Affairs Reference Handbook on Public-Sector AI. Most are sector-specific; the ones most relevant to “databases + AI” are the Copyright Act and the culture ministry’s generative-AI guidance.
Relevant laws
- The Artificial Intelligence Basic Act (AI Basic Act — passed Dec 2025, in force Jan 2026; the umbrella framework), the Personal Data Protection Act, the Copyright Act, the Trademark Act, the Patent Act, and the Trade Secrets Act.
Administrative guidance (cross-agency)
- MODA Administration for Digital Industries — Reference Guidelines for Evaluating AI Products and Systems (draft) (Mar 2024)
- NSTC — Guidelines on Using Generative AI in the Executive Yuan and its Agencies (Oct 2023)
- Taipei City Government — Guidelines on Using AI (Sep 2024)
Industry guidance (sector-specific)
- Ministry of Culture — Guidelines on Generative AI in Culture and the Arts (Jul 2025) — most relevant to creative work and copyright.
- Financial Supervisory Commission — Guidelines on AI in the Financial Industry (Jun 2024)
- Ministry of Health and Welfare — PCCP application guidance for AI/ML medical-device software (Sep 2024)
Policy white papers
- Executive Yuan — Chip-Driven Taiwan Industrial Innovation Program (Nov 2023)
- Financial Supervisory Commission — Core Principles and Policies for AI in Finance (Oct 2023)
- Executive Yuan — Taiwan AI Action Plan 2.0 (Feb 2023)
Want to go deeper? The rest is optional background.
Another question: who owns what AI creates?
So far we’ve talked about the input side (whether you can feed content to AI). This is the output side: whether what AI generates, translates, or rewrites for you can be protected, and who owns it. The answer varies by country, and it matters for whether you can claim any rights. Foreign rules are only for reference; in Taiwan, our Copyright Act and the TIPO’s views control.
Taiwan
- Made entirely by AI with no real human creative input → generally not protected by copyright.
- AI used as a tool, with real human creative input → the result can be protected, and the rights generally go to the person who created it.
- The level of human contribution is still decided case by case.
U.S.
- Made purely by AI, without real human authorship → generally not protected by copyright.
- Thaler v. Perlmutter: the Supreme Court declined to hear it in March 2026, keeping this rule.
Hong Kong / U.K. (common law)
- Computer-generated works can be protected — as long as enough skill, labor, or judgment went in.
- The author is “the person who made the necessary arrangements.” Protection lasts 50 years. This is broader than Taiwan or the U.S.
Mainland China
- In Nov 2023, the Beijing Internet Court found, in one case, that AI-generated content can be protected.
- This is a single case, not a general rule, and China is a separate legal system from Hong Kong and Taiwan.
The flip side: keeping AI from training on your own work
If you’re an author or a teacher and want to keep AI from scraping or training on your papers, slides, or other work, here are some options.
Images and artwork
Use Glaze (protective — it makes your style harder to copy) or Nightshade (a countermeasure — it turns images into “poison samples” that disrupt training). Both come from the University of Chicago.Glaze ↗ Nightshade ↗
Websites and text
Set up robots.txt to block crawlers. The ai.robots.txt project lists AI crawlers, and the EFF explains how.ai.robots.txt ↗ EFF ↗
Publishing contracts
When you sign a publishing contract, you can add the Authors Guild’s model “no AI training” clause (2025).Author's Guild ↗
Track licensing deals
Use Ithaka S+R’s Generative AI Licensing Agreement Tracker to see which publishers have signed AI deals.Ithaka Tracker ↗
How data owners help AI find and surface their content (open for details)
A different angle: what if you or your institution owns the content (a library, an archive, a publisher, an author)? How do you make sure AI can find and cite it? More and more people ask ChatGPT or Perplexity instead of searching, and each AI answer usually cites only a handful of sources — if yours is left out, you disappear from this new front door. Here are a few directions. (Database platforms’ built-in AI tools like Scopus AI and JSTOR are on a separate Library page → NCCU Library · database AI features ↗.)
Let AI crawl and cite you (GEO)
- Let AI crawlers in: allow GPTBot, PerplexityBot, ClaudeBot, and Google-Extended in robots.txt — this is the opposite of the “protect your work” section: if you want to be cited, they have to be able to read you.
- Publish an llms.txt: a file at your domain root — “a sitemap for AI” — listing your most citation-worthy content (about 70% of sites don’t have one yet).
- Make content extractable: answer first, headings phrased as real questions, one idea per paragraph, plenty of lists and tables, and a FAQ or TL;DR — so AI can grab and quote it.
- Make your identity clear: use structured data (schema) and align your institution/author entity with Wikipedia and Wikidata, so AI attributes you correctly.
Turn collections into open data (Collections as Data)
- Package collections as open datasets and APIs so AI can search and study them directly — the field calls this “collections as data.”
- Examples: Japan’s NDL (full-text AI search plus CC BY open datasets and OCR models), Harvard’s Institutional Books (394 million pages, AI-ready), HathiTrust’s full-text features dataset, the British Library’s collections as data, APIs from France’s Gallica / the Netherlands’ KB / Norway’s National Library / Europeana, the U.S. Library of Congress’s LC Labs, and the National Library of Korea’s AI training datasets.
- Taiwan: the National Central Library, NCCU and others keep digitizing and moving to open access, but turning collections into datasets AI can use directly is still at an early stage.
How publishers give content to AI (business models)
Model A: license your content to AI companies to “train” their models
- The publisher sells content to an AI company to train its model, usually for a one-time or installment license fee.
- Academic examples: Taylor & Francis × Microsoft (about US$10M), Springer Nature × Google (reported at about US$23M), and Wiley (about US$49M in AI licensing in fiscal 2026). Cambridge lets authors opt in instead.
- OpenAI is the most active buyer. Its chair, Bret Taylor, testified in 2026 that these licenses are “to avoid litigation.”
- The debate: most authors aren’t told and don’t get paid, so some argue for open access instead.
Model B: plug your content into an AI “answer engine,” shown live in answers with citations
- Your content shows up directly inside the AI’s answers, with source links, so users read it right there. This model is “live citation,” not using your content to train a model.
- Perplexity’s Publishers’ Program: every time the AI cites your content, you get a share of the ad or subscription revenue (publishers keep 80% under the Comet Plus plan). Unlike OpenAI’s one-off licensing fee, here the more you’re cited, the more you earn.
- Mistral × AFP works this way too (explicitly “not for training,” only live citation).
- Wiley plugged its STM (science, technology, medicine) materials into Perplexity’s answer engine.
- Students and staff can use the Wiley content their school already bought right inside Perplexity (in fields like nursing, business, and engineering). Answers include citations you can trace back.
- Access comes through the school’s Perplexity Enterprise Pro subscription. Wiley is Perplexity’s first education partner (announced May 2025). Pilot schools include Texas A&M and Texas State.
- For comparison: smaller publishers can’t negotiate deals on their own, so collective licensing is growing — for example, the U.S. Copyright Clearance Center (CCC) AI license and the U.K.’s CLA.
Newest: connecting databases straight to AI with MCP
A newer approach that started in late 2025. MCP (Model Context Protocol) is an open standard from Anthropic that OpenAI, Google, and Microsoft now support (it’s maintained by the Linux Foundation). Think of it as “a USB-C port for AI”: it lets tools like ChatGPT and Claude connect straight to a publisher’s server and pull peer-reviewed full text, with citations, in real time. That’s more direct and reliable than pasting a PDF yourself or asking the AI to search the web.
Who’s doing it, and how to connect
- Wiley launched its “Scholar Gateway / AI Gateway” MCP server (Oct 2025, with Anthropic). It connects about 3 million articles and works with Claude, Perplexity, and Mistral.
- Scite.ai has an MCP server (February 2026) covering 1.6 billion “smart citations.” Statista has one too. The platform provider Silverchair uses “Discovery Bridge” to bring more publishers online quickly.
- How to connect: the publisher runs an “MCP server.” You add it as a “connector” in Claude or ChatGPT, and the AI can then search and pull full text directly, with sources cited.
Does it cost money? What to watch for
- Cost: two models so far. Most work through your institution’s existing subscription (if your school subscribes, you can connect). Some charge per use (pay-per-event). It’s early, and pricing isn’t settled.
- Risks: the MCP server can see your queries. There are security concerns like prompt injection. Most are still in beta.
- Why it matters for libraries: which databases offer an official MCP, and how it’s licensed and priced, will become new questions for purchasing and patron services.
- You can connect OA yourself: the community has built many MCP servers that plug open content straight into Claude or ChatGPT, with no one else’s portal in between — for example OpenAlex (an open catalog of 240M+ scholarly works), paper-search-mcp (which federates dozens of OA sources at once: arXiv, PubMed, Unpaywall, CORE, Europe PMC, DOAJ, Zenodo, and more), and Zotero MCP for your own reference library.
- Paywalled content stays out: these reach only OA full text and open metadata. Full text behind a subscription wall (Elsevier, Wiley, and so on) can’t be reached by a self-hosted MCP — which is exactly why publishers’ official gateways exist.
- The red line: only OA is free to connect. Using an MCP in any way to get around a paywall and obtain unlicensed full text is the core dispute in Bartz and Elsevier v. Meta — it crosses the infringement line, so don’t do it.
- “Every library catalog” isn’t there yet: an open scholarly catalog (OpenAlex) and a personal library (Zotero) work today, but federating every library’s OPAC into one MCP is still at the experiment stage.
- Public-domain heritage: the cleanest use: no paywall, no licensing red line. The community has already turned Smithsonian Open Access (about 3 million CC0 items) and Wikidata into MCPs, and open sources like DPLA, Europeana, the Internet Archive, and Wikimedia Commons can be connected too. Institution-run official servers are still emerging; the model widely seen as best practice is to expose only metadata and send people back to the original — so the AI gives traceable, first-source pointers instead of ingesting the whole item, and the institution keeps control.
Glossary (open any term you don’t know)
- Text and data mining (TDM)
- Letting a computer read and analyze many texts at once, instead of reading them one by one.
- Fair use
- Copyright law lets you use someone’s work without permission in certain cases. Taiwan weighs four factors, case by case, so it’s never guaranteed. Important: it’s decided by a court case by case and is usually raised only when you’re accused of infringement — treat it as a defense, not a permission slip you can rely on in advance.
- Reproduction (copying)
- Making a copy of a work, including uploading it to a server. Whether it infringes depends on permission or fair use.
- Adaptation
- Re-creating a work into a new version, such as a translation or a rewrite. This right belongs only to the author.
- RAG (retrieval-augmented generation)
- The AI looks up sources first, then answers from them. Answers usually include citations, so it’s less likely to make things up.
- MCP (Model Context Protocol)
- An open standard that lets AI plug straight into a data source — like a USB-C port for AI.
- Opt-out
- When a rights holder declares in advance that their work may not be used by AI.
- Answer engine
- An AI search tool, like Perplexity, that gives you a direct answer with sources.
- Open access (OA) / CC license
- The author shares the work under a more open license. Still check which CC license it is (NC bans commercial use; ND bans changes).
- Public domain
- Works whose copyright has expired, so anyone can use them freely.
Full list of sources
Key court cases
Regional rules & libraries opening data (how the law regulates vs. how data is opened to AI)
Taiwan law and commentary
EU and Japan
Publisher deals with outside AI
Library AI, protecting your work, and tools
This is an information summary, not legal advice. The Library prepared it for outreach and education. It is not a lawyer’s opinion, and it does not guarantee or endorse whether any particular use is legal or counts as fair use. Copyright and fair use are decided case by case, and this field changes fast (checked Aug 28, 2026). For NCCU’s official policy or license purchases, confirm the exact contract with the Library, the Office of Research and Development, or the university’s legal office — and consult a lawyer if needed.