Reflections on Teaching Language AI for Research
I have been teaching the use of Language AI for Research, predominantly in humanities (including psychology and business administration) and legal studies, in different constellations now. It is still one of the subjects that I find most challenging; not only because of the fast pace of progress along the technical dimension, but also because supposedly local, technical questions easily evolve into fundamental discussions around research ethics, epistemology, or the goal of the entire research enterprise. In this post, I share some provisional insights on principles that have so far held up, and on what structures and approaches I found to work well.
Principle 0: Where to abstain
As a kind of preface: I do think that it is deeply problematic to enthusiastically encourage AI use in BA- and MA-degree courses. Acquiring the necessary intellectual and practical virtues, knowledge, and the deep skills (hopefully) characteristic of a university degree is best done with minimal reliance on AI tools. Everything else is akin to bringing the proverbial forklift to a gym (or an electronic mountain bike on a trail, for that matter): It confuses an activity where the production is the goal, as when you have to lift heavy loads efficiently in a storage facility, with an activity where formation of the actor is the goal, as in gym workouts or university degrees (or mountain biking).
In research, in contrast, the product does matter, arguably more than formation of the character of those involved (with PhD students, as so often, being somewhere in between). Therefore, it does not seem to be problematic as such to use AI, and Language AI in particular, when doing research.
So, principle 0 is: Before teaching language AI to students, first ask whether it is even appropriate to encourage language AI use to your specific target audience.
Principle 1: Don’t use AI to get faster, use it to get better
Mostly, AI is expected to help increase speed of research output (Chubb et al., 2022): “Use these tools to do more, and faster – so you’ll have a quantitative edge over your competitors”. I think this actually works, as the reviewing crisis in NLP shows: The quantity of prima facie legitimate papers has grown to a degree that the existing pool of reviewers cannot keep up. However, I think this emphasis on quantity is a result of a misunderstanding of the goal of research of the humanities and legal studies, together with inappropriate incentives especially for junior researchers. Likely, even on an individual level, that strategy will backfire: As AI keeps improving, the relevance of quantity will decrease – and the relevance of quality and genuine novelty will increase. So, the right strategy even from an individual perspective seems to prioritize quality over quantity. Needless to say, this is exactly what would solve the reviewing crisis.
Fortunately, AI can genuinely help improve the quality of research. It can give a more focused and comprehensive literature review, it can serve as a critical friend in experimental design, improve data analysis pipelines, etc. However, none of this will speed up the research process itself. Note the least because the LLMs are so far quite poor innovators, and there are good reasons to think that this is an inherent limitation of the technology: Devices optimized to predict the most probable next token are not natural born radical innovators. As a consequence, innovation is still a human job, and a hard, slow, and ironically often frustratingly boring one.
So, the first principle is: When teaching Language AI use in research, do not emphasize quantity or increased output, but emphasize and teach how it can help increase quality of output.
Principle 2: Integrate epistemological, ethical, and technical aspects
To me, this is the fascinating part of teaching Language AI for Research: A sustainable, satisfactory approach always has to involve research ethics as well as epistemological and technical considerations.
Epistemological
As a matter of fact, this is often overlooked, with the near-exclusive focus being placed on ethical and technical questions. I think this is unwise (and having a PhD in Philosophy with an emphasis on epistemology, this domain is naturally dear to me). Epistemology is the systematic study of knowledge: What knowledge is, what it means for a claim to be justified, what distinguishes good science from poor science. For instance, relying on hallucinated references or, in legal studies, hallucinated cases are primarily epistemological, not ethical failures.
Overall, the challenge is how to design AI-assisted workflows that honor the justificatory standards of the discipline. More specifically, often basic questions of epistemic autonomy and trust emerge: If a researcher cannot comprehend the code she uses to analyze the data anymore – or does not take the time to comprehend the code – then does she genuinely know what the results are? And how is this situation different from giving the task to a human assistant, e.g., a PhD student? One way you could think it is: We have evolved to be quite good at judging the reliability of fellow humans, while the same does not naturally hold for LLMs. This is compounded by the black-box nature of the technology as well as by the complete lack of transparency of frontier LLM providers.
Ethical
The ethical dimension is likely the most well-documented: Among others, it connects to data handling (is it permissible to send data of type X to a cloud-hosted LLM?), or experimental design (what do you need to tell participants in advance, even if it might threaten the overall set-up?). Also questions of academic citizenship emerge: For instance, beyond editorial guidelines, what are ethical considerations to take when using AI to assist in reviewing? Blind spots here can quickly turn into a PR disaster, see the case here. However, overall, my impression is that users are much more aware of these ethical pitfalls than of the epistemological ones, which is why I often emphasize the epistemological dimension a bit more.
Technological
The technical aspect is the hands-on craft, and it has to be basically reassessed every quarter: it involves prompting strategies (even more so in the age of agentic AI), tool comparison for specific tasks, systematic verification loops, etc. What matters most is not that scholars leave with a list of tools — those will change — but that they leave with the judgement to assess emerging tools on their own.
How do these three dimensions interact? Imagine a new tool that enters the market, which is much more capable of visualizing data, and which is also explainable. For both reasons, it would score highly on the epistemological dimension. However, imagine as well that it requires sending sensible data to servers with questionable data protection laws, creating an ethical problem. This might then be a reason to go back on the market and look for an alternative tool which is maybe not optimal on the epistemological dimension (less explainable, for instance), but which does provide sufficient data security.
So, the second principle is: Always integrate these three dimensions when teaching language AI for research.
What this means for teaching
On my experience, there are roughly three levels – and course durations – of teaching the subject.
The first is orientation: about a day to establish what contemporary language AI is and isn’t, accompanied by an exploration of the ethical and epistemological commitments any scholarly use must respect. This is the minimum responsible exposure — enough to leave with a working frame and avoid the most embarrassing mistakes, not enough to leave with practised methods.
The second is practice: another day on top, this time hands-on. Prompting strategies, retrieval for primary sources, output evaluation, a guided exercise on participants’ own material. Theory becomes practical method; tools get tested against real research, participants feel the complexity of the situation on their own skin, with the three dimensions in isolation often recommending different tools or procedures.
The third is embedded work: a further day that takes a specific project and designs an AI-assisted workflow around it — literature, data, writing, peer-review preparation, ethical documentation, technical setup. What comes out the other side is a pipeline tailored to a particular research context, rather than a generic methodology to be ported and lost.
Why I am prioritizing this
From what I observe in the NLP community and also in philosophy publishing, it seems that language AI is the greatest challenge to reasonable, epistemically sound academic research in my generation. It could be very beneficial, leading to much better research, but what I am currently seeing is worrying. Researchers use language AI to increase quantity, not quality of research, sometimes barely understanding their results (or English, which is the language of their publication). So, if you have any suggestions or advice on this matter, I would be grateful if you could reach out. Of course, if you think I could contribute something from my side, I would also be very happy to hear about it, feel free to contact me, perhaps most easily by email.
Bibliography
2022
Enjoy Reading This Article?
Here are some more articles you might like to read next: