, , , , , , ,

UK’s AI Safety Institute easily jailbreaks major LLMs

In a shocking turn of events, AI systems might not be as safe as their creators make them out to be — who saw that coming, right? In a new report, the UK government's AI Safety Institute (AISI) found that the four undisclosed LLMs tested were "highly vulnerable to basic jailbreaks." Some unjailbroken models even generated "harmful outputs" without researchers attempting to produce them.

Most publicly available LLMs have certain safeguards built in to prevent them from generating harmful or illegal responses; jailbreaking simply means tricking the model into ignoring those safeguards. AISI did this using prompts from a recent standardized evaluation framework as well as prompts it developed in-house. The models all responded to at least a few harmful questions even without a jailbreak attempt. Once AISI attempted "relatively simple attacks" though, all responded to between 98 and 100 percent of harmful questions.

UK Prime Minister Rishi Sunak announced plans to open the AISI at the end of October 2023, and it launched on November 2. It's meant to "carefully test new types of frontier AI before and after they are released to address the potentially harmful capabilities of AI models, including exploring all the risks, from social harms like bias and misinformation to the most unlikely but extreme risk, such as humanity losing control of AI completely."

The AISI's report indicates that whatever safety measures these LLMs currently deploy are insufficient. The Institute plans to complete further testing on other AI models, and is developing more evaluations and metrics for each area of concern.

This article originally appeared on Engadget at https://www.engadget.com/uks-ai-safety-institute-easily-jailbreaks-major-llms-133903699.html?src=rss

https://www.engadget.com/uks-ai-safety-institute-easily-jailbreaks-major-llms-133903699.html?src=rss


July 2024
M T W T F S S
1234567
891011121314
15161718192021
22232425262728
293031  

About Us

Welcome to encircle News! We are a cutting-edge technology news company that is dedicated to bringing you the latest and greatest in everything tech. From automobiles to drones, software to hardware, we’ve got you covered.

At encircle News, we believe that technology is more than just a tool, it’s a way of life. And we’re here to help you stay on top of all the latest trends and developments in this ever-evolving field. We know that technology is constantly changing, and that can be overwhelming, but we’re here to make it easy for you to keep up.

We’re a team of tech enthusiasts who are passionate about everything tech and love to share our knowledge with others. We believe that technology should be accessible to everyone, and we’re here to make sure it is. Our mission is to provide you with fun, engaging, and informative content that helps you to understand and embrace the latest technologies.

From the newest cars on the road to the latest drones taking to the skies, we’ve got you covered. We also dive deep into the world of software and hardware, bringing you the latest updates on everything from operating systems to processors.

So whether you’re a tech enthusiast, a business professional, or just someone who wants to stay up-to-date on the latest advancements in technology, encircle News is the place for you. Join us on this exciting journey and be a part of shaping the future.

Podcasts

TWiT 986: Our Dope GPS! – Supreme Court Decisions, Snapdragon X Elite Tests This Week in Tech (Audio)

Supreme Court Decisions, Snapdragon X Elite Tests Murthy Decision Should Not Foreclose Cases Against Actual First Amendment Violations What SCOTUS just did to broadband, the right to repair, the environment, and more Nearly 4,000 arrested in global police crackdown on online scam networks Mark Cuban's public email was hacked after receiving call from a fake Google rep The Julian Assange Saga Is Finally Over Microsoft's bundling of Office and Teams breaks antitrust law, EU says EU Competition Commissioner says Apple's decision to pull AI from EU shows anticompetitive behavior Microsoft says it's okay to steal content published on the web Microsoft's Surface Laptop 7 Copilot+ PC is finally the best clamshell laptop on the market after 8 years of iterations Tested: Don't buy a Snapdragon X Elite laptop for PC gaming Signal 65 Snapdragon battery testing The RIAA's lawsuit against generative music startups will be the bloodbath AI needs Wherein The Copia Institute Asks The Second Circuit To Stand Up For Fair Use, The Internet Archive, And Why We Bother To Have Copyright Law At All Redbox's owner files for bankruptcy after repeatedly missing payments and payroll Host: Leo Laporte Guests: Cathy Gellis, Ryan Shrout, and Doc Rock Download or subscribe to this show at https://twit.tv/shows/this-week-in-tech Get episodes ad-free with Club TWiT at https://twit.tv/clubtwit Sponsors: lookout.com 1password.com/twit
  1. TWiT 986: Our Dope GPS! – Supreme Court Decisions, Snapdragon X Elite Tests
  2. TWiT 985: TikTok With Wings – AT&T Landlines, US Bans Kaspersky and DJI
  3. TWiT 984: Fifty-three Clicks – Bot Farms in Ukraine, LA Public Health Dept. Phished
  4. TWiT 983: Digital Snackwells – NVIDIA's Thor, Adobe's TOS, Insta's Unskippable Ads
  5. TWiT 982: International Trash – Startup Chaos, Breaking Ticketmaster, Ultrasonic Coffee