Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    The Download: tricking LLMs, and reviving geothermal plants

    Your iPhone will soon have a room for a driver's license, as long as you live in Oklahoma and these other five states

    A fundamental flaw leaves LLMs strikingly vulnerable to attack

    Facebook Twitter Instagram
    • Tech
    • Gadgets
    • Spotlight
    • Gaming
    Facebook Twitter Instagram
    circuitthoughtscircuitthoughts
    Subscribe
    • Home
    • Gadgets
    • Insights
    • Apps

      Google Uses AI Searches To Detect If Someone Is In Crisis

      Gboard Magic Wand Button Will Covert Your Text To Emojis

      Android 10 & Older Devices Now Getting Automatic App Permissions Reset

      Spotify Blend Update Increases Group Sizes, Adds Celebrity Blends

      Samsung May Improve Battery Significantly With Galaxy Watch 5

    • Gear
    • Mobiles
      1. Tech
      2. Gadgets
      3. Insights
      4. View All

      The Download: tricking LLMs, and reviving geothermal plants

      Your iPhone will soon have a room for a driver's license, as long as you live in Oklahoma and these other five states

      A fundamental flaw leaves LLMs strikingly vulnerable to attack

      Razr (2025) drops to just $549.99, perfect for those who don’t want to overspend

      March Update May Have Weakened The Haptics For Pixel 6 Users

      Project 'Diamond' Is The Galaxy S23, Not A Rollable Smartphone

      The At A Glance Widget Is More Useful After March Update

      Pre-Order The OnePlus 10 Pro For Just $1 In The US

      Motorola Edge+ Review: It Checks A Lot Of Boxes

      This Smartphone Concept Design Is Different… In A Good Way

      Twitter Just Made Searching Your Direct Messages Better

      That Netflix Price Hike Is Starting To Take Place

      Latest Huawei Mobiles P50 and P50 Pro Feature Kirin Chips

      Samsung Galaxy M62 Benchmarked with Galaxy Note10’s Chipset

      9.1

      Review: T-Mobile Winning 5G Race Around the World

      8.9

      Samsung Galaxy S21 Ultra Review: the New King of Android Phones

    • Computing
    circuitthoughtscircuitthoughts
    Home»Tech»A fundamental flaw leaves LLMs strikingly vulnerable to attack
    Tech

    A fundamental flaw leaves LLMs strikingly vulnerable to attack

    adminBy No Comments3 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email

    “When you and I are talking, I can tell which words are coming out of my mouth because I can feel my mouth moving,” says Cui. But an LLM just sees a continuous stream of text; a user’s prompts are mixed up with the model’s previous responses, scratch-pad notes, text copied from documents, and so on. “It’s just one big sheet of tokens,” she says.

    To help keep track of who said what, chatbots use tags to break the text up by what researchers call roles. Everything you type gets put between tags, and everything the LLM writes back gets put between tags. Text provided by a model’s designers to guide its core behavior is put between tags, text that a model generates in its chain of thought is put between tags, and text that a model picks up from an external source, such as a web page or another agent, gets put between tags. (Cui says that these are the labels OpenAI uses for its models; other firms might use different ones. The purpose is the same, however.)

    Roles have become the foundation on which LLMs are trained to resist hacks, because most attacks boil down to tricking the model into acting as if an instruction came from someone or something it did not. For example, many jailbreaks (where a user tricks a model into saying or doing things its makers do not want it to) work by making a model read text as if it were or text. And many prompt injections (where a hacker slips a model new instructions) work by making a model read text as if it were , , or text.

    When model makers train LLMs to resist attacks, a lot of it comes down to getting the models to spot when instructions pop up in places they shouldn’t.  

    But what Cui and her colleagues discovered is that LLMs are in fact very bad at keeping track of different roles. In a series of experiments that looked at what was going on inside a handful of different models, the researchers found that LLMs seem to identify the role of a specific chunk of text not by the tags around it but by the style of that text and the words it contains.

    They found that swapping tags around—replacing tags with tags, for example—made almost no difference to how the LLM interpreted the text itself. If it looked like text from its own chain of thought, then the LLM acted as if it really were. Ditto for all other roles.  

    Weak link

    The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem.

    “I like this paper a lot,” says Florian Tramèr, a computer scientist who works on LLMs and cybersecurity at ETH Zürich. The attack insight is really neat, he says.

    #fundamental #flaw #leaves #LLMs #strikingly #vulnerable #attack

    attack flaw fundamental leaves LLMs strikingly vulnerable
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

    Related Posts

    The Download: tricking LLMs, and reviving geothermal plants

    Your iPhone will soon have a room for a driver's license, as long as you live in Oklahoma and these other five states

    Razr (2025) drops to just $549.99, perfect for those who don’t want to overspend

    Add A Comment

    Leave A Reply Cancel Reply

    Editors Picks
    8.5

    Apple Planning Big Mac Redesign and Half-Sized Old Mac

    Autonomous Driving Startup Attracts Chinese Investor

    Onboard Cameras Allow Disabled Quadcopters to Fly

    Top Reviews
    9.1

    Review: T-Mobile Winning 5G Race Around the World

    By
    8.9

    Samsung Galaxy S21 Ultra Review: the New King of Android Phones

    By
    8.9

    Xiaomi Mi 10: New Variant with Snapdragon 870 Review

    By
    circuitthoughts
    Facebook Twitter Instagram Pinterest Vimeo YouTube
    • Home
    • Tech
    • Gadgets
    • Mobiles
    • Our Authors
    © 2026 ThemeSphere. Designed by WPfastworld.

    Type above and press Enter to search. Press Esc to cancel.