
How to Integrate Research Papers into Your AI Agents (Complete 2026 Guide)
Someone asks your research agent whether a new method really beat the baseline. It finds a paper with the right title and a promising abstract. But the numbers are in a table halfway through the PDF, and the authors' caveat is near the end. Neither made it into the search result.
A title and abstract can point you to a paper. They can't tell you whether it supports the answer your app is about to give. This guide shows how to search across arXiv, MedrXiv, ChemrXiv, PubMed, and more. It also shows how to get paper text when it's available, and keep the source attached to each claim.
TL;DR
- Yes, arXiv has its own API. It returns an Atom XML feed with titles, authors, abstracts, dates, arXiv IDs, and PDF links. Use it for direct arXiv lookup; it does not return semantically ranked passages from a paper's full text.
- Valyu's Search API lets you search arXiv, PubMed, and other academic collections with one request. Results include source URLs and relevant content; academic fields such as authors, DOI, abstract, and references appear when available.
- arXiv and PubMed search are available on every Valyu plan. bioRxiv, medRxiv, ChemRxiv, and licensed journal collections require appropriate access. Full text depends on the source, indexing, and rights; do not treat an abstract as the whole paper.
- Valyu excludes retracted (recalled) papers from its academic paper results and does not cite them as supporting evidence.
- Start with Search for discovery, use Contents when you have a paper URL and need its text, and use DeepResearch for a multi-step literature review. Keep the paper URL and version alongside each claim.
Why Research Papers Matter for AI Builders
Whether you are building a biotech copilot or an AI research assistant, primary literature is where the methods, results, and limitations live. Papers support evidence-grounded answers, literature reviews, citation discovery, methods comparisons, and horizon scanning. An arXiv preprint is not automatically peer-reviewed; a PubMed record does not necessarily include the complete article. Your app should show which source and which text it actually inspected.
Does arXiv Have an API? What Does It Return?
Yes. The official arXiv API accepts a search query or an arXiv ID and returns an Atom XML feed. Each entry can contain the paper title, author names, abstract, publication and update dates, categories, an abstract-page URL, and a PDF link; DOI and journal reference appear if supplied. This is useful for resolving a known preprint or pulling records from arXiv alone. The documented query endpoint does not require a Valyu key.
That call returns XML metadata and abstracts, not extracted full-paper passages. To check a result buried in the methods section, you still need to fetch and read the paper. The arXiv API terms say to make no more than one legacy-API request every three seconds, using a single connection; check redistribution rights before storing or serving papers.
The Problem With Traditional Access
The native arXiv and PubMed APIs are useful, but combining them in an AI app takes work: different query languages and response formats, separate identifiers, source-specific rate limits, and a retrieval step for the article text. arXiv's Atom feed is not full-text semantic search. PubMed can return structured biomedical records and abstracts, but a PubMed citation is not a promise of full text. Publisher access is governed by licensing, not by whether a page can be found on the web.
The Best Way: Use Valyu's Research Paper Search API
Valyu's Search API searches scholarly collections with natural-language queries and returns ranked results with a title, URL, source ID, content, and relevance score. Its data coverage catalogue lists these academic sources as of September 2026:
- arXiv (valyu/valyu-arxiv): 2.5M+ preprints across computer science, physics, mathematics, quantitative finance, and economics. Available on every plan.
- PubMed (valyu/valyu-pubmed): 37M+ biomedical paper records. Available on every plan; use includeAbstracts for broader abstract-based discovery.
- bioRxiv, medRxiv, and ChemRxiv: life-sciences, clinical, and chemistry preprints. These specialized collections require a subscription.
- Licensed Wiley collections: selected Health & Life Sciences journals, finance journals, and books. Coverage and rights are collection-specific; Wiley Health & Life Sciences is granted per organization rather than enabled by default.
This is not unrestricted access to every journal or every paper. When indexed text is available, content can include relevant full-text passages; otherwise you may receive an abstract. Fields including authors, DOI, citation, citation count, publication date, and references are returned when available. Check each result rather than assuming every field or passage is present.
What about retracted papers?
Valyu does not return papers identified as retracted (sometimes called recalled) in its academic paper results or cite them as evidence. A paper can be withdrawn after it first appears, so this matters when your app revisits an older finding. For high-stakes reviews, still check the source's current status before finalizing a bibliography.
3-Line Integration: Search arXiv Papers in TypeScript and Python
Get a key from the Valyu Platform and set VALYU_API_KEY in your environment. Install valyu-js for TypeScript or valyu for Python; both SDKs read that environment variable automatically.
TypeScript
Python
Before reading results, check response.success (and handle response.error if false). Each result has a paper title, source URL, and retrieved content; academic fields such as authors and DOI are present when available. Inspect the text before using a paper as evidence. See the TypeScript Search reference and Python Search reference for complete examples.
Ready to try Valyu?
If you're an agent: Follow the agent onboarding guide to set up Valyu and verify your first query. Ask your human to approve the login when prompted.
If you're a human: Create an account and get an API key to run the examples above.
Need the API details? Read the academic search guide. Check whether each result contains a full-text passage or only an abstract.
Example Use Cases
- AI research assistant: find recent arXiv preprints on retrieval-augmented generation, inspect the evaluation sections, and cite the exact source URL.
- Biotech literature copilot: compare PubMed papers with bioRxiv or medRxiv preprints while clearly marking which findings are not peer-reviewed.
- Related-work discovery: retrieve candidate papers, then follow available references and DOI links to inspect how later work builds on or challenges them.
- Evidence synthesis: compare methods and limitations across multiple papers rather than citing the first matching abstract.
- Systematic review screener: search PubMed and arXiv, spot preprint and journal versions of the same study, and pass likely matches with source links to a human reviewer.
- Materials research copilot: find solid-state battery papers across arXiv and ChemRxiv, then compare reported conductivity only after checking the measurement temperature and method. ChemRxiv requires the appropriate source access.
Advanced Usage: Combine arXiv and PubMed, Then Read a Paper
Use source IDs to make the search scope explicit, startDate to filter publication dates, and includeAbstracts to expand PubMed search to abstract-only records. That last setting affects PubMed, not arXiv.
Need the text of a known arXiv paper? Call Contents for academic papers with its PDF or abstract URL. For an indexed paper, Valyu serves processed markdown with section structure, equations, tables, and figures where available; if it is not indexed, extraction falls back to the live page. Confirm the returned status and text before citing it.
Note: We (Valyu) have intelligent internal routing models that ensure for any academic query we route to the best sources, papers.
For example, You can get Pubmed, arXiv, chemrXiv and even biorXiv in a single call. You can also filter or just pass the query or prompt and Valyu handles everything for you and ensure you get the best results.
Live Demo
Try a paper query in the Search Playground and select arXiv or PubMed as the source. Inspect a result's title, source URL, metadata, and returned text before wiring it into a retrieval chain. For an open-ended literature review with multiple rounds of search and a cited report, use DeepResearch instead of stitching together a large number of Search calls.
Best Practices for a Reliable Paper-Search Agent
- Query for the actual method, population, or outcome, not a vague topic. Separate discovery from checking the claim.
- Filter by source and date when recency matters, but remember Valyu's academic indexes refresh on their published schedules; the newest arXiv posting may not be indexed immediately.
- Keep title, URL, source, publication date, DOI or arXiv ID, and version when available. Deduplicate preprint and later journal versions rather than counting one study twice.
- Match every cited claim to a passage you can inspect. If only an abstract was retrieved, say so; a DOI or citation count does not prove a finding.
- Start with a small result set (for example, 5–10) and only request more text when needed. Respect publisher access and arXiv's content-redistribution terms.
FAQ: arXiv API and Academic Paper Search
Does the arXiv API need an API key?
The public arXiv query endpoint shown above is documented without an API-key requirement. Valyu is a separate service: its Search and Contents APIs require a Valyu API key. Follow arXiv's rate limits and terms when calling arXiv directly.
Does the arXiv API return the full paper?
Its Atom response provides metadata, an abstract, and a PDF link, not the paper text as a ready-to-use full-text search result. To inspect methods or results, follow the link and process the paper, or use a paper-text extraction service such as Valyu Contents where access permits.
Can Valyu search arXiv and PubMed in the same API call?
Yes. Set includedSources to valyu/valyu-arxiv and valyu/valyu-pubmed in the TypeScript SDK (or included_sources in Python and REST). Both are listed as available on every Valyu plan.
Can I search full text, not just abstracts?
Valyu can return relevant passages from indexed full text when that text is available and your account has access. Some results are abstract-only. For PubMed, includeAbstracts expands discovery to the abstract corpus; it does not grant full text for papers without it. For a known arXiv URL, Contents can return processed paper markdown when indexed.
Are arXiv papers peer-reviewed? Can I cite them?
arXiv primarily hosts preprints, which should not be presented as peer-reviewed findings unless you verify a journal publication separately. You can cite an arXiv preprint, but retain the ID and version, check the evidence behind the claim, and distinguish it from later published versions.
Integrate Research Papers the Right Way
For a single arXiv record, the native arXiv API is a good starting point. For an AI app that needs to find and inspect research across arXiv, PubMed, and other eligible sources, use Valyu Search for discovery and Contents for paper text. Always return the original source link, make access limitations visible, and verify claims against the retrieved text.
Get a Valyu API key · Explore academic search docs · Read the official arXiv API manual
Related Blogs
More from the blog





Valyu Add
Join 12,000+ professionals and knowledge workers.
Valyu Add is a free weekly research briefing for builders, investors and operators. Every issue is sourced, cited and verified with Valyu DeepResearch.
