{"id":1332,"date":"2026-08-07T14:09:56","date_gmt":"2026-08-07T14:09:56","guid":{"rendered":"https:\/\/redzine.co.uk\/index.php\/2026\/08\/07\/ai-is-making-disinformation-harder-to-spot-but-weve-found-a-new-way-to-catch-it\/"},"modified":"2026-08-07T14:09:56","modified_gmt":"2026-08-07T14:09:56","slug":"ai-is-making-disinformation-harder-to-spot-but-weve-found-a-new-way-to-catch-it","status":"publish","type":"post","link":"https:\/\/redzine.co.uk\/index.php\/2026\/08\/07\/ai-is-making-disinformation-harder-to-spot-but-weve-found-a-new-way-to-catch-it\/","title":{"rendered":"AI is making disinformation harder to spot \u2013 but we\u2019ve found a new way to catch it"},"content":{"rendered":"<p>Do you ever see comments on social media that seem way off topic, but still manage to wrench the discussion around to divisive political debate?<\/p>\n<p>A discussion about the cost of living suddenly becomes an argument about immigration. A conversation about the war in Ukraine turns into claims about government corruption. It can feel jarring \u2013 and sometimes this is deliberate.<\/p>\n<p>As generative AI becomes more powerful, malicious groups are increasingly <a href=\"https:\/\/theconversation.com\/how-we-tricked-ai-chatbots-into-creating-misinformation-despite-safety-measures-264184\">using it<\/a> to produce and spread <a href=\"https:\/\/theconversation.com\/topics\/disinformation-42353\">disinformation<\/a> online. Automated accounts can flood social media with convincing comments designed to sow <a href=\"https:\/\/www.bbc.co.uk\/news\/articles\/cp3l4kwer5ko\">division<\/a>, inflame political debate and undermine trust in reliable information.<\/p>\n<p>But our <a href=\"https:\/\/doi.org\/10.1016\/j.dcm.2025.100929\">latest research<\/a> offers a way to spot these attempts. Rather than trying to identify whether a post was written by AI, we focus on something different: whether it\u2019s trying to derail the conversation.<\/p>\n<p>Until recently, identifying malicious accounts was often quite straightforward. Many campaigns relied on people writing in a second language. So, posts sometimes contained grammatical mistakes or unusual word choices. Detection systems could look for these patterns in the language used.<\/p>\n<p>But generative AI has changed that. AI systems can now produce fluent, natural-sounding text that is much harder to distinguish from human writing. For example, patterns like use of <a href=\"https:\/\/www.merriam-webster.com\/grammar\/em-dash-en-dash-how-to-use\">em-dashes<\/a> and the word \u201cdelve\u201d used to be <a href=\"https:\/\/theconversation.com\/too-many-em-dashes-weird-words-like-delves-spotting-text-written-by-chatgpt-is-still-more-art-than-science-259629\">telltale signs<\/a> of a text being generated by AI. But AIs are adapting, and these older systems are <a href=\"https:\/\/theconversation.com\/why-its-so-hard-to-tell-if-a-piece-of-text-was-written-by-ai-even-for-ai-265181\">increasingly ineffective<\/a>. <\/p>\n<p>Trying to detect AI purely from the words people use is becoming a losing battle. We believe the better approach is to look at what a message is trying to achieve. <\/p>\n<h2>Looking for signs<\/h2>\n<p>Attempts to spread disinformation often work by steering conversations away from their original topic, towards more polarising issues. So, instead of analysing individual words, we set out to build a system that could recognise this phenomenon in online discussions.<\/p>\n<p>We analysed comments posted beneath BBC News videos on YouTube, a platform that has previously been <a href=\"https:\/\/doi.org\/10.1177\/1940161220912682\">targeted<\/a> by organised disinformation campaigns.<\/p>\n<p>For example, imagine a comment about Ukraine\u2019s president, Volodymyr Zelensky, interacting with senior UK political figures: \u201cZelensky must be wondering how many foreign secretaries the UK goes through.\u201d Now, imagine another person responding: \u201cMind you, Zelensky has barely been president for four years. Maybe that\u2019s why the little tyrant bans his opposition.\u201d <\/p>\n<p>Whether that second point is true or false is not the issue. Instead of responding to the original comment, it redirects the conversation towards a different, more divisive topic.  <\/p>\n<p>This is known as a red herring: introducing an unrelated issue that distracts from the original discussion. These kinds of shift are difficult for conventional disinformation detection systems to identify, because they are not tied to particular words or phrases.<\/p>\n<figure class=\"align-center \">\n            <img decoding=\"async\" alt=\"Composite photo collage of a person sitting at a chair with a manipulative hand hanging over them.\" src=\"https:\/\/images.theconversation.com\/files\/752262\/original\/file-20260805-50-z8yqrt.jpg?ixlib=rb-4.1.1&amp;q=45&amp;auto=format&amp;w=754&amp;fit=clip\"><figcaption>\n              <span class=\"caption\">36% of derailing messages online included \u2018red herrings\u2019.<\/span><br \/>\n              <span class=\"attribution\"><a class=\"source\" href=\"https:\/\/www.shutterstock.com\/image-photo\/composite-photo-collage-anonym-girl-sit-2519338975\">Roman Samborskyi\/Shutterstock<\/a><\/span><br \/>\n            <\/figcaption><\/figure>\n<p>We manually analysed more than 1,600 comments under BBC News videos, labelling them according to 25 different features of online discussion \u2013 and discovered some clear patterns. <\/p>\n<p>We found that 36% of derailing messages had red herrings, 65% had leaps in logic known as \u201cnon sequiturs\u201d, and 20% contained personal attacks. They were also much less likely to acknowledge previous comments or express empathy.<\/p>\n<h2>Spotting manipulation<\/h2>\n<p>The next step was to see whether an AI system could recognise these patterns automatically. We used an AI to catch an AI. <\/p>\n<p>For every genuine online comment, we asked an AI large language model to generate several reasonable, relevant responses. Returning to the example of UK foreign secretaries, the AI suggested replies such as: \u201cThe current situation in this country must come as quite a shock\u201d or \u201cOne too many?\u201d. Both responded directly to the original point. <\/p>\n<p>The system then compares the real response with our AI-generated replies. If the actual comment differs substantially, it may indicate that someone is attempting to steer the conversation in a different direction. So, rather than searching for suspicious words, our system looks for unexpected changes in the flow of the discussion.<\/p>\n<p><strong>How the system works:<\/strong><\/p>\n<figure class=\"align-center \">\n            <img decoding=\"async\" alt=\"A diagram of a system for detecting derailing discourse\" src=\"https:\/\/images.theconversation.com\/files\/752098\/original\/file-20260804-50-ew3nsk.png?ixlib=rb-4.1.1&amp;q=45&amp;auto=format&amp;w=754&amp;fit=clip\"><figcaption>\n              <span class=\"caption\">Discourse derailment is measured by the distance between the real reply and a set of expected replies generated by an AI.<\/span><br \/>\n              <span class=\"attribution\"><a class=\"source\" href=\"https:\/\/doi.org\/10.1057\/s41599-026-07900-x\">Krykoniuk, Hopkin-King &amp; Roberts: Using LLMs to identify discourse derailment as a potential cue for disinformation in social media posts (2026).<\/a>, <a class=\"license\" href=\"http:\/\/creativecommons.org\/licenses\/by\/4.0\/\">CC BY<\/a><\/span><br \/>\n            <\/figcaption><\/figure>\n<p>We tested this approach using our manually labelled dataset. In <a href=\"https:\/\/doi.org\/10.1057\/s41599-026-07900-x\">our second study<\/a>, the system correctly identified derailing comments around 77% of the time.<\/p>\n<p>That\u2019s far from perfect, but no detection system is \u2013 particularly when analysing something as complex as human conversation. However, our approach performed around twice as well as existing systems based on word-level sentiment analysis. It also achieved results comparable with the level of agreement between human researchers.<\/p>\n<p>Our approach is effective because the AI learns what a typical response to a conversation looks like. When a reply unexpectedly changes the discussion, the system can identify that change and analyse patterns that earlier methods couldn\u2019t detect.<\/p>\n<hr>\n<p>\n  <em><br \/>\n    <strong><br \/>\n      Read more:<br \/>\n      <a href=\"https:\/\/theconversation.com\/why-science-gcses-matter-more-than-we-think-in-a-post-truth-age-276306\">Why science GCSEs matter more than we think in a post-truth age<\/a><br \/>\n    <\/strong><br \/>\n  <\/em>\n<\/p>\n<hr>\n<p>Of course, going off topic isn\u2019t necessarily a sign of malicious intent or disinformation. People naturally take conversations in unexpected directions, and there are many legitimate reasons why discussions evolve.<\/p>\n<p>For that reason, this technology may act as an early-warning system rather than a replacement for human judgment. It could help moderators identify conversations that deserve closer attention \u2013 but any final decisions should remain with trained experts.<\/p>\n<p>There are also important ethical questions to address. AI systems can reflect biases in the data they are trained on, and they still do not understand conversations in quite the same way that people do. Improving how AI represents and interprets human discussion remains a challenge.<\/p>\n<p>As AI-generated content becomes increasingly difficult to distinguish from human writing, detecting disinformation requires more than simply searching for telltale words. It requires understanding how conversations work, how they are manipulated, and when someone is trying to quietly steer them off course.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/counter.theconversation.com\/content\/288671\/count.gif\" alt=\"The Conversation\" width=\"1\" height=\"1\" \/><\/p>\n<p class=\"fine-print\"><em><span>Se\u00e1n Roberts was funded by the AI and autonomy for intelligence, surveillance and reconnaissance (A2ISR) project at the Defence Science and Technology Laboratory (Dstl) through the Defence and Security Accelerator (DASA), which is part of the UK Ministry of Defence, United Kingdom.<\/span><\/em><\/p>\n<p class=\"fine-print\"><em><span>Kateryna Krykoniuk was funded by the AI and autonomy for intelligence, surveillance and reconnaissance (A2ISR) project at the Defence Science and Technology Laboratory (Dstl) through the Defence and Security Accelerator (DASA), which is part of the UK Ministry of Defence, United Kingdom.<\/span><\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Do you ever see comments on social media that seem way off topic, but still manage to wrench the discussion around to divisive political debate? A discussion about the cost of living suddenly becomes an argument about immigration. A conversation about the war in Ukraine turns into claims about government corruption. It can feel jarring [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1332","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/posts\/1332","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/comments?post=1332"}],"version-history":[{"count":0,"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/posts\/1332\/revisions"}],"wp:attachment":[{"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/media?parent=1332"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/categories?post=1332"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/redzine.co.uk\/index.php\/wp-json\/wp\/v2\/tags?post=1332"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}