- Models: https://weteachwithai.com/models/
- Desktop Agents
- Duke Devils’ Advocate: https://oit.duke.edu/devils-advocate/
- System Prompts: https://weteachwithai.com/system-prompts/
- API Tools: https://weteachwithai.com/apis/
AI RESEARCH WORKFLOW
1. Move to Desktop (1× 10-minute set-up)
A chatbot can only tell you things. An agent can read your folder, run code on your data, open a database you are logged into, write a file, and come back with what it could not do.
- Install a desktop agent. Claude Cowork (claude.ai/download), the ChatGPT desktop app, Microsoft 365 Copilot, or Gemini in Workspace. They differ in detail; the workflow below is the same in all of them.
- Make a Project per research project. A Project is a persistent container: instructions, files and memory that carry across sessions. Put your standing rules in it — evidence standards, citation style, what counts as a source in your field. Write these once and then things start to happen automatically.
- Connect a folder. Give the agent a real folder on your computer. Unless you are building a database with the AI, you can keep raw data or original files in one subfolder and let the agent write only to another. I’ve never had a problem, but you can add a rule like always create a new file to your project instructions.
- Connect a browser. A browser extension (Claude in Chrome) or agentic browser (Comet, Edge Copilot Mode) lets the agent work inside databases you are already logged into — library proxies, digitized archives, catalogues that block ordinary scrapers.
2. Literature Review
You can use either template (below) or a third-party tool like Consensus (above) to do your literature review, OR you can do it inside the project. Either way, you will want to create a document with your references and share it with your AI project.
LIMITS: Fabricated citations are measurably increasing in the published record (although the audit was AI-assisted!) Science citations are easier to verify, so assume the humanities rate might be higher and harder to detect. Using a third-party or a good search prompt (see above) can dramatically reduce fabricated citations, but a human should verify. The tools below are still required to provide further checks.
- Topaz, M., et al. (2026, May 7). Fabricated citations: an audit across 2.5 million biomedical papers. The Lancet. Link
- Orrall, A. (2026, May 7). One in 277 PubMed-indexed papers in 2026 shows fabricated references, says analysis. Retraction Watch. Link
Research Review Template 1
You now have as many research assistants as you want, but just like humans, you will need to guide them. You can do a lot with a single detailed prompt:
- Create a research report that will illuminate/examine/explore X. Make sure to examine the questions A, B, and C and include an analysis of D & E. You should begin with a critical review of literature/practice/web and then provide a synthesis of the key ideas/controversies/concepts/case studies and a recommendation.
- Sources & Scope: The research should draw from fields F & G; methodology H; focus on peer-reviewed journal articles/best practices/reputable studies/institutional sources; look for sector/Western/political/educational/gender bias in sources; seek global sources in language/culture I.
- Purpose & Framework: Use K as a framework for understanding these issues. Focus on real-world applications and capabilities. Pay special attention to policy implications and government uses. Note any potential for L.
- Audience: Write for an audience of M/for journal N or submission to conference O. Describe your findings with relevance to P.
OR use Semantic Scholar RAGS: Consensus, Elicit, Undermind, Research Rabbit, or Edison Scientific — an autonomous “AI scientist” for long, multi-step literature analyses (edisonscientific.com).
Research Review Template 2
An iterative sequence can sometimes be better than one long prompt. Pasting each prompt in twice (before you hit return) also seems to improve results.
- Literature Search. [You can often use the prompt above or one of the API tools below to do this within the Semantic Scholar database of published academic papers. Providing the actual papers (or links) improves the quality of what comes next.]
- Map the Landscape. Organize this list of papers. Group them into clusters of shared assumptions, claims, methodologies and/or data sets. Create a table. List the papers in column 1 and then list the core claim (in 50 words or less) in column 2. In column 3 list the key methodology or assumption that guides this paper. In column 4 list all ideas in that paper that are contradicted by other papers and cite those papers.
- Big Idea Lineage. List the central claims and/or the most contentious issues or methods in this literature. First create a table that includes: (a) by whom and in what paper the idea was introduced, (b) who are/were the primary challengers, (c) summarize the positions on either side, (d) explain why they disagree and (e) tell me if there is now any consensus. Then also create a structured knowledge map or family tree of this literature that shows how these ideas have interacted.
- Mine the Gaps. Based on all of these papers and this analysis, identify 10 big research questions that are still unanswered. Describe the gaps and why they exist. Cite the papers which have come closest. What assumptions do most of these papers share, but do not explicitly justify or test? State these assumptions and cite a few of the important papers that rely on them most. What useful data or method is most underused?
- Summarize. Briefly summarize in less than 500 words what the field believes collectively, what is proven beyond a reasonable doubt, what remains contested and what is the single most important unanswered question. What would happen to the field if its most important assumption turned out to be wrong?
- Getting Started. Explain all of this in 300 words to a non-expert without jargon. Summarize what is known, what is unknown and where this matters in the real world. List the 3 most important papers I should read first to get a grip on this field.
Copy, paste & modify a Deep Research Template at WeTeachWithAI.com/workshop. Or ask an AI to write a prompt for you.
3. Harvest Data
AI agents are good at repetitive tasks but they are literal and you need to create clear rules. Start a new chat and first have a conversation about what is available.
- My goal is X. I would like to build a database where I can do Y. Make a complete list of available sources and anticipate problems and where you might face difficulty gaining access to the best available data.
- Have a conversation, push back and solve problems. (You might need to log in to your library site to order ILL or get access to journals, for example.) Then write a prompt that will both collect good data and make sure AI tells you the truth about what it could not find. It is especially useful to use two different models to check each other. I asked Fable to write harvest handoffs for Sonnet in batches small enough to be accurate. Fable then checked every batch and wrote a new handoff (47 in total).
Harvest Prompt Template
- HARVEST [items] from [named archives, catalogues, databases]. For each item record [fields A-D]. RULES: one row per item; record the source URL and the date you retrieved it for every field; if a field is unknown, leave it empty — never infer, never fill from general knowledge; flag duplicates and identity collisions (ex. same name, different person) rather than merging them; list what you found and excluded, and why. OUTPUT a CSV plus a short memo of everything you could not verify.
- Then, before you use any of it, run this skew check and REPORT the composition of what you have collected [by language, country, decade, source etc.]. Which of these is over-represented, and is that a finding or an artifact of what happens to be digitized? What should the next pass target to correct it?
Example: AI Literacy Programs
I want to make a comprehensive list of AI literacy programs, microcredentials, certificates and badges offered at universities anywhere in the world. Here are some examples… Think about the categories—name of the program, name of institution, format (online, self-paced, credit), length, required or optional (and if so required of what group and when), does it come with a digital badge or certificate? First do the comprehensive survey and collect the information and then create a spreadsheet. Result: view the spreadsheet.
Example: Handoff Prompt
- HANDOFF <Project> Session <N>: <one-line focus> (<author>, <date>)
- FIND artifacts <Governing doc / scope reminder.>
- SAVE <current_file_or_version>. Save any new version as <next_version>, refuse-if-exists.
- WRITE what this session did: what changed and the key number; a method note if non-obvious; the single most important result or unlock, stated plainly.
- REPORT what and how things were verified: how you checked the work — test, count reconciliation, spot-check, independent re-derivation. State what you did NOT verify.
- DO, in priority order: the highest-value next step (with why it matters and the concrete first move); the second (with status: started/blocked/untouched); the third (lower priority, one line).
- KNOWN limitations / disclosures: where the work is precise-but-not-exhaustive; a decision made under uncertainty and the reasoning, so it isn’t silently treated as fact.
- CRIB these reusable techniques.
LIMITS: Anticipate corpus bias. Training data drawn from Common Crawl and similar sources is heavily skewed toward English (even when you instruct it to do a global search) and toward recently digitized web text. You as the human need to be ever vigilant. Vargas-Parada, L. (2025, Nov 27). Large language models are biased — local initiatives are fighting for change. Link
4. BS Testing
Depending on your discipline and project there are many options here. It is often better to do each individually. (They can be part of the same chat but should be individual prompts.) Even for humanities work.
4a. Data Analysis & Skew Check
Ask for something a summary cannot fake. Insist on an audit trail and on testing the tool against itself.
ADJUST BY DISCIPLINE: Which repositories are digitized and which were not? Is provenance data important? What to do with empty fields? Are there identity collisions (same name, different person)? Are quotations being checked against the right edition? Start in another language? Sampling frame?
Analysis Prompt Template. Review the [corpus/dataset/set of texts]. Do not summarize it.
- (1) Describe the shape of this data set: what is actually being counted, what is missing, what the sampling frame really is.
- (2) Give me the three patterns you would defend to a hostile referee — each with the specific rows or passages that support it, and the strongest reason each might be an artifact of the data rather than a fact about the world.
- (3) Give me one pattern you noticed that you cannot explain. Rate your confidence in each and tell me which parts I must check by hand.
- (4) Apply my codebook/rules/guidelines [paste it] to these [interviews / field notes / documents]. Record the exact passage for every code applied. Do not invent codes — list passages that fit no existing code separately as candidates. Then run the whole codebook a second time and show me where the two runs disagree.
LIMITS: Since the same prompt on the same text can produce different results, run important things twice. Disagreement between runs is a reliability estimate. Log and keep your prompts, model, version, and date. Set temperature to 0 (no randomness) where the tool allows it. Disclosing AI use in your methods section is an obligation (at least in Europe).European Commission (2026, May 8). Living guidelines on the responsible use of generative AI in research (revised). Link. Main, J. B. (2026). Thematic analysis with open-source generative AI and machine learning: a new method for inductive qualitative codebook development. Humanities and Social Sciences Communications, 13, 209. Link
4b. Verify Citations Template
- Check every reference: resolve the DOI or permalink and confirm authors, title, journal, year and pages against the publisher record. Mark every reference (or create a new document) VERIFIED / WRONG DETAIL / NOT FOUND. Then check each verified source for retractions, expressions of concern, published critiques, or failed replications. Do not correct anything silently — show me each discrepancy and let me decide. For any source you cannot verify through a publisher of record, tell me where it was actually published and what review it passed. If a paper’s own provenance is doubtful, say so plainly and tell me not to cite its numbers — I may still want it as an example of the problem.
4c. Check Claims Support Template
- Your goal is to find where I am wrong: counter-evidence to my existing claims is the most valuable thing you can produce. Examine each and every claim in my draft and the source attached to it. Does the cited source/tradition/theory actually support this specific claim, at this strength? Comment (using the comment function) on any claim that is overstated, unsupported or uncited. If there is a citation, verify it and mark any faulty citations. If there is a better or newer citation, put it in the comments. Always trace citations to the original study, never to the press release, the news write-up, or a secondhand citation. Where I have overstated, quote the sentence from the source that shows the real strength of the finding. If an assertion does not have a citation, but should, then find the best citation and list it in the comments.
- FOLLOW-UP: I have now edited and revised this document. Add all of the new and revised citations to the bibliography.
4d. Newer, Better, or Contradictory Citations Template
- For each of my central claims and citations, search for: (i) work published since my citation that supersedes or qualifies it; (ii) the strongest published evidence against the claim; (iii) failed replications, null results, and publication-bias corrections in this literature; (iv) relevant work in languages other than English and from outside [region]. Report the contradictions first, before anything that agrees with me.
- Rate every source using this rubric or evidence ladder [depending on your field: secondary source, randomized trial, systematic review or meta-analysis, quasi-experiment, correlational survey, preprint, vendor telemetry, practitioner report etc]. Investigate sample size, setting, author reputation etc. Flag any famous-but-fragile source whose citation count exceeds its evidentiary weight. Be vigilant about flagging any low-quality evidence.
4e. How Wrong Could I Be? Template
- Here is my bibliography. Report what this evidence base is made of — by method, discipline, education level, country, language of publication, and year. What is over-represented? Which of my claims rest on a part of the literature that is thinner than the claim implies?
- Draw a random sample of X [quotes, data points etc.] and hand-check them against the primary source, showing your work for each. Report the error rate in your own verification, the kinds of errors you made, and which categories of item you are least reliable on.
- Never bulk-apply a match on an ambiguous identifier. Shared surnames, variant spellings and namesakes get hand-checked individually, never flipped in a batch. One confident bulk edit can corrupt a dataset in ways that are very hard to find later.
- When something cannot be verified, the correct output is a blank and a note — not a plausible guess. A guessed value is indistinguishable from a checked one once it is in the file.
Example: Design Principles for Teaching in the Age of AI
https://weteachwithai.com/designing-learning-with-ai/
- Labeled both a widely circulated EEG study and the most-cited paper in the field as carrying enormous public weight relative to their designs.
- Found a bias-adjusted reanalysis showing no effect.
- Highlighted that the studies are geographically broader than usual, but overwhelmingly quantitative and experimental, and concentrated in STEM and language learning. There is very little qualitative classroom research and almost nothing from arts and humanities teaching.
5. Organize and Update
Agents are careless with your files unless told otherwise. Make safety checks part of your instructions.
- Insert these references into [file] in alphabetical order as tracked changes. Skip any that are already present. Count the references before and after and confirm the arithmetic. Validate that the file still opens. If anything does not match, abort and report rather than saving.
- Write a handoff brief for the next session: what was done; what is staged but not yet merged; what is unresolved and why; the two highest-value next tasks; and the mistakes not to repeat.
Example: Book Project
Your goal is to set me up to revise X for Y. Act like a senior editor with deep experience in Z. You will not be writing any text. We have a word limit of # including citations. Here is the publisher style guide.
- Read A, B & C [abstract, proposal, first version, ToC, notes etc.].
- Create a new folder D and make new copies (so new doc files) of the files in E. Label your new files Book F ch# etc. Keep the same formatting from the style sheet. Make new versions of references, introduction etc. as well.
- Add a new file for G based on the ToC and use the other chapters for style but then use the contents H.
- Create a complete references doc and include complete citations for everything cited in the book in format I.
- Now read each chapter and COMMENT only on: (a) any passage that needs to be updated or cut; (b) any passage or citation where the evidence is weak, providing a needed citation in the comments and adding it to the references; (c) any better or newer citation, also in the comments; (d) keep track of the references as you read and compare them to the references document, adding items in the text but missing from the references, and flagging items in the references doc that are NOT in the text (with another color highlight—don’t delete them).
- When done, create a new ToC document with your annotations and recommendations for additional updates, chapters or sections.
- You can use folder J [with article pdfs, notes and your research] as needed or do your own research. I am especially interested in K. My audience is L. Put it all into the ToC document v4_annotations.
- Create a new folder on Google Docs and upload all of this inside that folder with comments tracked.
6. Feedback
- What might an average reader/college professor/IRS auditor find confusing/objectionable/exciting?
- Give me feedback from a reader who misunderstands my intentions.
- Create feedback that will challenge me; include feedback with inaccurate information.
- Analyze this to help me improve it. Respond as…
- You are an experienced expert in X and also a senior editor at Y. Analyze this and help me re-imagine, improve and transform it using Z.
Example: A Panel of Scholars
- Do multiple individual analyses of this text by a wide group of scholars with different views of the field. Start with X and Y [pick your own list of scholars/rivals] and include recent work on Z. Include scholars from these competing theoretical areas [list] or the authors of [list of books].
- Then have each one read all of the other reviews and compare what they think is good and what needs improvement. Create a dashboard chart that lists the recommended changes and how many and which advisors recommend each. Cluster the recommendations agreed by most advisors at the top. Write a brief report that gives me three solid suggestions for improvement and also highlights what parts of this research are most contested. Provide citations for the work that most challenges this. Use this analysis to ask me thought-provoking questions about how I might improve my chances of acceptance at the Journal of X.
Save Reusable Skills
When you have refined a set of rules for a task you will repeat, save it as a Skill so you never have to re-explain it. Examples:
- Evaluate the credibility of sources in this way…
- Apply my rubric criterion by criterion; feedback first, grade last; coach by question, not correction.
- Extract a text’s thesis, premises and hidden assumptions first; steelman before critiquing.
- Sort evidence in this way…
- Apply my codebook to interviews or field notes in this way…
Example: Case Chronology
List all of the evidence for this case in chronological order. List the source of each time point. Highlight all disagreements explicitly. Your rule: put what is firmly established and uncontested in bold.
Example: Research Updates (Revised by Claude for itself)
- When I share a new study or report, your job is to verify it, integrate it into the manuscript apparatus, and keep all copies in sync. Accuracy is paramount: never add anything unverified, always distinguish evidence tiers, and hold positions only when the evidence supports them.
- Where everything lives: Manuscript (source of truth): [Directory]. Drive mirror: [Google link]. Slide citations: [file names]. Website: [web address and browser].
- Step A — Verify before anything else. Fetch the actual source (web_fetch the DOI/SSRN/arXiv/publisher page or PDF; use WebSearch if only a citation fragment is given). Confirm authors, year, exact title, venue, and the specific numbers/claims to be cited — quote figures from the source, not from press coverage. Check for retractions, corrections, and published critiques. Classify the evidence tier and carry it into every note you write:
- RCTs and randomized field experiments;
- Systematic reviews / meta-analyses;
- Quasi-experiments and large-N administrative data;
- Surveys and perception data (attitudes, never outcomes);
- Preprints / working papers;
- Practitioner / press sources (cite as practice, not evidence).
- House rule: famous-but-fragile studies (contested preprints, corrected papers) are cited only as cultural phenomena, never as evidence. If a new source contradicts something already in the manuscript, that is the most valuable outcome. If verification fails or is ambiguous, report that and stop — do not integrate.
- Step B — Integrate into the manuscript. Add the APA entry to the references doc as a tracked insertion, alphabetically, deduped, attributed to “Claude”. Write one comment per chapter, labeled ADD / UPDATE / CUT / COUNTER-EVIDENCE / EVIDENCE CHECK, with the finding, its key numbers, the evidence tier, and where it fits. The safety protocol — never skip it: count existing comments, run the script, verify the count equals before + n, validate, and only then save (never delete on the mount). If any check fails, ABORT that file and say so.
- Step C — Slide-citation documents. If the finding is one Jose might cite in workshops, append the full citation to the Workshop Slides and Citations.docx, matching each document’s existing format. Do not touch any .pptx, and leave the other citation documents alone unless asked.
- Step D — Website. Via Claude in Chrome: navigate to the edit URL, find the appropriate section, and add the citation in the page’s existing format. Two rules: (1) preview the exact text to Jose in chat before saving, and (2) clicking Update publishes to a public site — get his explicit OK in chat first, every time.
- Reporting back. Close with a compact summary: what was verified (and its tier), which chapters got notes and why, what went into references/slides/website, what is synced vs. pending, and any caveat. Brief prose, no ceremony.
Vibe Coding, Websites and Visualizations
1. Create a Simulation
- Play: https://josebowen.github.io/BinghamGame/
- See more about vibe coding here: https://weteachwithai.com/vibe-coding/.
- See more about simulations here: https://weteachwithai.com/simulations-games-roleplaying/
2. Schedule a Task
Every Monday at 7am, review all of my course data in the LMS (chats, completion, scores, time-on-task etc.) and compare it to the baseline averages from previous weeks. Flag anything that’s shifted more than 10% from baseline. Send me an email organized by course with a flag for any metric that has changed and include any possible explanation and remedies. Tell me how I can better help my students.
3. Create an Interactive Database and Website
Example: Women Piano Composers
See the full prompt, workflow and the versions produced by different AI models: https://weteachwithai.com/vibe-coding/ — and the result: https://josebowen.github.io/women-at-the-keys/
- Start by doing an exhaustive search of published academic work for resources about 19th century women composers (born from 1800 to 1900) who wrote classical piano music. Look for books, articles, archives, encyclopedias, indexes and websites. Examples include A Guide to Piano Music by Women Composers, vol 1 by Pamela Youngdahl Dees. Sources are often incomplete, so find multiple sources.
- Create a searchable database of 19th century female composers of piano music, including any women who published verifiable sheet music, born between 1800 and 1900. This should be a global index but include all European countries. Search Wikipedia, IMSLP, rare piano scores and the sources discovered in step 1.
- Then make a nicely-formatted list organized by country and, within that, a chronological list of women composers born in that country. Include the key and duration of each piece and a link to sheet music or archives. List only solo piano compositions. Use this format: COUNTRY → Last name, First name (birth and death dates) → brief biographical information → Piano composition opus 1 (duration, key, links to sheet music) → opus 2, etc.
- Create the code (and html file) for a website that provides searchable access to this database. An excellent model is the Swedish Musical Heritage. Searchable elements should include country, composer name, title, key and piece duration.
- Create and deploy a public website where I can share this searchable database. Ask me any clarifying questions before beginning.
- Order ILLs: look through my pdf library and identify pieces I am missing that are not available as pdf downloads from IMSLP, find archival sources on WorldCat, and order them for me from SMU Interlibrary Loans.
Example: Piano Repertoire Explorer
I’ve always been interested in the 19th century piano repertoire beyond Beethoven, Brahms and the boys. To understand better what composers and works were popular in different cities and how repertoire moved around, I needed AI. You can now interact with over 250,000 published piano works with a new (v1) piano repertoire explorer: https://josebowen.github.io/piano-repertoire-explorer/
- After working on this for decades, I spent a week (!?!) having Fable and Sonnet collect publication records, catalogues, reviews, and other info into one huge data set.
- Then it deduplicated, cross-referenced, checked and tagged the data at scale. (Joe Verdi and G. Verdi are probably the same composer if associated with an Italian opera, but only solo piano arrangements of overtures should be counted.)
- Then Fable built a range of different tools and visualizations to explore this data.
AI can scale things far beyond my human limits. If I do manage to write a new history of piano repertoire in the 19th century, it will only be because AI helped me.
Will AI Make Your Research More Productive?
- …productivity among GenAI users rose by 15% in 2023 relative to non-users and further increased to 36% in 2024… mean impact factors rising by 1.3 percent in 2023 and 2.0 percent in 2024. Filimonović, D., Rutzer, C., & Wunsch, C. (2025, Oct). Can GenAI Improve Academic Performance? arxiv.org/abs/2510.02408 (p. 14)
- “Pre-Read” tools for journals and conferences: refine.ink and paperreview.ai.
- AI “accelerates manuscript output, reduces barriers for non-native English speakers, and diversifies the discovery of prior literatures,” but “traditional signals of scientific quality such as language complexity are becoming unreliable indicators of merit.” Kusumegi, K., et al. (2025). Science 390, 1240–1243. DOI:10.1126/science.adw3000