There are multiple guidelines from publishers and organizations on the use of artificial intelligence (AI) in publishing.1–5 However, none are specific to family medicine. Most journals have some basic AI use recommendations for authors, but more explicit direction is needed, as not all AI tools are the same.
As family medicine journal editors, we want to provide a unified statement about AI in academic publishing for authors, editors, publishers, and peer reviewers based on our current understanding of the field. The technology is advancing rapidly. While text generated from early large language models (LLMs) was relatively easy to identify, text generated from newer versions is getting progressively better at imitating human language and more challenging to detect. Our goal is to develop a unified framework for managing AI in family medicine journals. As this is a rapidly evolving environment, we acknowledge that any such framework will need to continue to evolve. However, we also feel it is important to provide some guidance for where we are today.
Definitions: Artificial intelligence is a broad field where computers perform tasks that have historically been thought to require human intelligence. LLMs are a recent breakthrough in AI that allow computers to generate text that seems like it comes from a human. LLMs deal with language generation, while the broader term generative AI can also include AI generated images or figures. ChatGPT is one of the earliest and most widely used LLM models, but other companies have developed similar products. LLMs “learn” to do a multifaceted analysis of word sequences in a massive text training database and generate new sequences of words using a complex probability model. The model has a random component, so responses to the exact same prompt submitted multiple times will not be identical. LLMs can generate text that looks like a medical journal article in response to a prompt, but the article's content may or may not be accurate. LLMs may “confabulate,” generating convincing text that includes false information.6,7,8 LLMs do not search the internet for answers to questions. However, they have been paired with search engines in increasingly sophisticated ways. For the rest of this editorial, we will use the broad term AI synonymously with LLMs.
ROLE OF LARGE LANGUAGE MODELS IN ACADEMIC WRITING AND RESEARCH
As LLM tools are updated and authors and researchers become familiar with them, they will undoubtedly become more functional in assisting the research and writing process by improving efficiency and consistency. However, current research on the best use of these tools in publication is still lacking. A systematic review exploring the role of ChatGPT in literature searches found that most articles on the topic are commentaries, blog posts, and editorials, with little peer-reviewed research.9 Some studies have demonstrated benefit in narrowing the scope of literature review when AI tools were applied to large data sets of studies and prompted to evaluate them for inclusion based on the title and abstract. Another paper reported that AI had 70% accuracy in appropriately identifying relevant studies compared with human researchers and may reduce time and provide a less subjective approach to literature review.10,11,12 When used to assist with writing background sections, LLMs' writing was rated the same if not better than human researchers, but the citations were consistently false in another study.13 LLM models are frequently deficient in providing “real” papers and correctly matching authors to their own papers when generating citations and therefore are at risk of creating fictitious citations that appear convincing despite incorrect information including digital object identifier (DOI) numbers.6,14,
Read the full article
Get immediate access, anytime, anywhere.
Choose a single article, issue, or full-access subscription.
Earn up to 5 CME credits per issue.
