{"id":162,"date":"2026-09-17T06:59:17","date_gmt":"2026-09-17T06:59:17","guid":{"rendered":"https:\/\/www.graveiensai.com\/blog\/?p=162"},"modified":"2026-09-17T06:59:17","modified_gmt":"2026-09-17T06:59:17","slug":"multilingual-llm-red-teaming","status":"publish","type":"post","link":"https:\/\/www.graveiensai.com\/blog\/multilingual-llm-red-teaming\/","title":{"rendered":"Multilingual LLM Red Teaming: Why Safe in English Doesn\u2019t Mean Safe in Every Language"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Mul\u00adti\u00adlin\u00adgual LLM red team\u00ading is the prac\u00adtice of adver\u00adsar\u00adi\u00adal\u00adly test\u00ading a large lan\u00adguage mod\u00adel for unsafe behav\u00adior in every lan\u00adguage it serves, not only in Eng\u00adlish, using native speak\u00aders who write and grade orig\u00adi\u00adnal attack prompts in each tar\u00adget lan\u00adguage. It exists because a safe\u00adguard that holds in Eng\u00adlish often col\u00adlaps\u00ades when the same request is rephrased in Hin\u00addi, Ben\u00adgali, Tamil, Swahili, or Ara\u00adbic. Peer-reviewed work pre\u00adsent\u00aded at ICLR 2024 found that large lan\u00adguage mod\u00adels were rough\u00adly three times more like\u00adly to pro\u00adduce harm\u00adful con\u00adtent in low-resource lan\u00adguages than in high-resource ones, and that trans\u00adlat\u00ading an unsafe Eng\u00adlish prompt into a rarely test\u00aded lan\u00adguage could push unsafe-out\u00adput rates from sin\u00adgle dig\u00adits into the major\u00adi\u00adty of respons\u00ades. If your mod\u00adel ships in more than one lan\u00adguage, Eng\u00adlish-only test\u00ading is not a mea\u00adsure of its safe\u00adty. It is a mea\u00adsure of one lan\u00adguage\u2019s safe\u00adty.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>At a glance<\/strong><\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Ques\u00adtion<\/strong><\/th><th><strong>Short answer<\/strong><\/th><\/tr><\/thead><tbody><tr><td>What is mul\u00adti\u00adlin\u00adgual LLM red team\u00ading?<\/td><td>Adver\u00adsar\u00adi\u00adal safe\u00adty test\u00ading of an LLM across every lan\u00adguage it serves, using native speak\u00aders who author and grade attack prompts in each lan\u00adguage.<\/td><\/tr><tr><td>Why does it mat\u00adter?<\/td><td>Safe\u00adty train\u00ading is con\u00adcen\u00adtrat\u00aded in Eng\u00adlish, so mod\u00adels fail more often in oth\u00ader lan\u00adguages. Research shows low-resource lan\u00adguages face about 3x the harm\u00adful-out\u00adput rate of Eng\u00adlish.<\/td><\/tr><tr><td>What does it test for?<\/td><td>Jail\u00adbreaks and adver\u00adsar\u00adi\u00adal attacks, mul\u00adti\u00adlin\u00adgual tox\u00adi\u00adc\u00adi\u00adty and harm\u00adful-con\u00adtent eval\u00adu\u00ada\u00adtion, and mul\u00adti\u00adlin\u00adgual bias and fair\u00adness test\u00ading, across sin\u00adgle-turn and mul\u00adti-turn con\u00adver\u00adsa\u00adtions.<\/td><\/tr><tr><td>Why not machine-trans\u00adlate Eng\u00adlish probes?<\/td><td>Trans\u00adla\u00adtion miss\u00ades idioms, translit\u00ader\u00ada\u00adtion, and code-switch\u00ading, the exact pat\u00adterns real users and attack\u00aders use, so it under-reports true risk.<\/td><\/tr><tr><td>Who needs it?<\/td><td>AI lab safe\u00adty teams, third-par\u00adty eval\u00adu\u00ada\u00adtors and AI Safe\u00adty Insti\u00adtutes, and enter\u00adpris\u00ades deploy\u00ading mod\u00adels in mul\u00adti\u00adlin\u00adgual mar\u00adkets.<\/td><\/tr><tr><td>How is scope decid\u00aded?<\/td><td>By pri\u00ador\u00adi\u00adtiz\u00ading lan\u00adguages on expo\u00adsure, harm sever\u00adi\u00adty, lin\u00adguis\u00adtic com\u00adplex\u00adi\u00adty, and safe\u00adty-data scarci\u00adty, then test\u00ading across lan\u00adguage tiers, attack class\u00ades, and turn depth.<\/td><\/tr><tr><td>Does it make a mod\u00adel safe?<\/td><td>No. It pro\u00adduces evi\u00addence of where a mod\u00adel fails. Fix\u00ading those fail\u00adures through align\u00adment and guardrails is a sep\u00ada\u00adrate step.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Table of contents<\/strong><\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>What is mul\u00adti\u00adlin\u00adgual LLM red team\u00ading?<\/li>\n\n\n\n<li>Why Eng\u00adlish-only safe\u00adty test\u00ading breaks in oth\u00ader lan\u00adguages<\/li>\n\n\n\n<li>The main types of mul\u00adti\u00adlin\u00adgual safe\u00adty test\u00ading<\/li>\n\n\n\n<li>How mul\u00adti\u00adlin\u00adgual LLM red team\u00ading works<\/li>\n\n\n\n<li>The Graveiens LENS Frame\u00adwork for lan\u00adguage pri\u00ador\u00adi\u00adti\u00adza\u00adtion<\/li>\n\n\n\n<li>Approach\u00ades com\u00adpared: which method wins when<\/li>\n\n\n\n<li>The mul\u00adti\u00adlin\u00adgual safe\u00adty test\u00ading matu\u00adri\u00adty mod\u00adel<\/li>\n\n\n\n<li>What mul\u00adti\u00adlin\u00adgual red team\u00ading costs<\/li>\n\n\n\n<li>Com\u00admon mis\u00adtakes and how to avoid them<\/li>\n\n\n\n<li>A prac\u00adti\u00adcal mul\u00adti\u00adlin\u00adgual red-team\u00ading check\u00adlist<\/li>\n\n\n\n<li>Worked exam\u00adples<\/li>\n\n\n\n<li>How reg\u00adu\u00adla\u00adtion is rais\u00ading the bar<\/li>\n\n\n\n<li>Fre\u00adquent\u00adly asked ques\u00adtions<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>What is multilingual LLM red teaming?<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Red team\u00ading, in an AI con\u00adtext, is struc\u00adtured adver\u00adsar\u00adi\u00adal test\u00ading: skilled peo\u00adple delib\u00ader\u00adate\u00adly attack a sys\u00adtem with inputs designed to make it fail, then doc\u00adu\u00adment exact\u00adly where and how it breaks. Applied to a <a href=\"https:\/\/www.graveiensai.com\/blog\/what-is-an-llm\">large lan\u00adguage mod\u00adel<\/a>, red team\u00ading means craft\u00ading prompts that try to bypass the mod\u00adel\u2019s safe\u00adty guardrails so that fail\u00adures are found by a friend\u00adly team before they are found by users or bad actors.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Mul\u00adti\u00adlin\u00adgual LLM red team\u00ading extends that dis\u00adci\u00adpline across lan\u00adguages. Instead of assum\u00ading that an Eng\u00adlish safe\u00adty result gen\u00ader\u00adal\u00adizes, native speak\u00aders write orig\u00adi\u00adnal adver\u00adsar\u00adi\u00adal and jail\u00adbreak prompts in each tar\u00adget lan\u00adguage, sub\u00admit them to the mod\u00adel, and grade the respons\u00ades against a defined harm tax\u00adon\u00ado\u00admy. The out\u00adput is not a pass or fail badge. It is a labeled dataset and a per-lan\u00adguage find\u00adings report that shows which harms slip through in which lan\u00adguages, at what sever\u00adi\u00adty. That evi\u00addence then feeds eval\u00adu\u00ada\u00adtion dash\u00adboards and align\u00adment train\u00ading.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The dis\u00adtinc\u00adtion that mat\u00adters most: mul\u00adti\u00adlin\u00adgual red team\u00ading is about cov\u00ader\u00adage, not trans\u00adla\u00adtion. A mod\u00adel can be gen\u00aduine\u00adly safe in Eng\u00adlish and qui\u00adet\u00adly unsafe in a dozen oth\u00ader lan\u00adguages at the same time, and only lan\u00adguage-native test\u00ading reveals that gap.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Why English-only safety testing breaks in other languages<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The core rea\u00adson is data. Safe\u00adty align\u00adment, the fine-tun\u00ading that teach\u00ades a mod\u00adel to refuse harm\u00adful requests, is trained over\u00adwhelm\u00ading\u00adly on Eng\u00adlish exam\u00adples. Lan\u00adguages with less text on the inter\u00adnet receive less safe\u00adty train\u00ading, so their guardrails are thin\u00adner. Johns Hop\u00adkins researchers put it plain\u00adly: the root issue is that there sim\u00adply is not enough data avail\u00adable for less wide\u00adly used lan\u00adguages dur\u00ading a mod\u00adel\u2019s first train\u00ading process, which leaves safe\u00adty behav\u00adior under\u00adde\u00advel\u00adoped exact\u00adly where it is hard\u00adest to audit.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The mea\u00adsured effects are large and con\u00adsis\u00adtent across inde\u00adpen\u00addent stud\u00adies:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Low-resource lan\u00adguages car\u00adry far more risk.<\/strong> The ICLR 2024 study behind the Mul\u00adti\u00adJail bench\u00admark, built from 3,150 sam\u00adples across nine lan\u00adguages, found low-resource lan\u00adguages pro\u00adduced unsafe con\u00adtent about three times as often as high-resource lan\u00adguages when users were not even try\u00ading to attack the mod\u00adel. When a mali\u00adcious instruc\u00adtion was com\u00adbined with a low-resource lan\u00adguage, unsafe-out\u00adput rates for one wide\u00adly used mod\u00adel rose to rough\u00adly 80 per\u00adcent, and an adap\u00adtive attack approached near\u00adly 100 per\u00adcent.<\/li>\n\n\n\n<li><strong>Tox\u00adi\u00adc\u00adi\u00adty ris\u00ades as lan\u00adguage resources fall.<\/strong> Poly\u00adglo\u00adTox\u00adi\u00adc\u00adi\u00adtyPrompts, a bench\u00admark of 425,000 nat\u00adu\u00adral\u00adly occur\u00adring prompts across 17 lan\u00adguages eval\u00adu\u00adat\u00aded on 62 mod\u00adels, found that tox\u00adi\u00adc\u00adi\u00adty decreas\u00ades as the avail\u00adabil\u00adi\u00adty of lan\u00adguage resources increas\u00ades, describ\u00ading a per\u00adsis\u00adtent gap in mul\u00adti\u00adlin\u00adgual tox\u00adi\u00adc\u00adi\u00adty mit\u00adi\u00adga\u00adtion even in high\u00adly capa\u00adble mod\u00adels.<\/li>\n\n\n\n<li><strong>Con\u00adver\u00adsa\u00adtions and non-Latin scripts com\u00adpound the prob\u00adlem.<\/strong> Ama\u00adzon Sci\u00adence\u2019s mul\u00adti-turn, mul\u00adti\u00adlin\u00adgual red-team\u00ading work found mod\u00adels were on aver\u00adage 71 per\u00adcent more vul\u00adner\u00ada\u00adble after a five-turn Eng\u00adlish con\u00adver\u00adsa\u00adtion than after a sin\u00adgle turn, and that non-Eng\u00adlish, non-Latin-script lan\u00adguages reached a 68 per\u00adcent aver\u00adage attack suc\u00adcess rate ver\u00adsus about 41 per\u00adcent for Eng\u00adlish.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Both mul\u00adti\u00adlin\u00adgual tox\u00adi\u00adc\u00adi\u00adty and harm\u00adful-con\u00adtent eval\u00adu\u00ada\u00adtion and mul\u00adti\u00adlin\u00adgual bias and fair\u00adness test\u00ading show the same pat\u00adtern: results degrade as lan\u00adguage resources fall, so a mod\u00adel that looks clean in Eng\u00adlish can car\u00adry mea\u00adsur\u00adable tox\u00adi\u00adc\u00adi\u00adty and skewed treat\u00adment in oth\u00ader lan\u00adguages. Machine trans\u00adla\u00adtion does not res\u00adcue an Eng\u00adlish probe set. Real attack\u00aders and real users mix scripts, translit\u00ader\u00adate, and code-switch mid-sen\u00adtence, and those pat\u00adterns are pre\u00adcise\u00adly what a lit\u00ader\u00adal trans\u00adla\u00adtion flat\u00adtens out. That is why seri\u00adous pro\u00adgrams pair lan\u00adguage-native probe design with grad\u00aded, ratio\u00adnale-rich <a href=\"https:\/\/www.graveiensai.com\/llm-evaluation\">LLM eval\u00adu\u00ada\u00adtion<\/a> rather than treat\u00ading a trans\u00adlat\u00aded test as mul\u00adti\u00adlin\u00adgual cov\u00ader\u00adage.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Also read: <a href=\"https:\/\/www.graveiensai.com\/blog\/what-is-rlhf\">What Is RLHF?<\/a> for how grad\u00aded human feed\u00adback turns red-team find\u00adings into align\u00adment train\u00ading data.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>The main types of multilingual safety testing<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Mul\u00adti\u00adlin\u00adgual red team\u00ading is an umbrel\u00adla over sev\u00ader\u00adal dis\u00adtinct test\u00ading types, most impor\u00adtant\u00adly jail\u00adbreak prob\u00ading, mul\u00adti\u00adlin\u00adgual tox\u00adi\u00adc\u00adi\u00adty and harm\u00adful-con\u00adtent eval\u00adu\u00ada\u00adtion, and mul\u00adti\u00adlin\u00adgual bias and fair\u00adness test\u00ading. Strong pro\u00adgrams run all of them, because each sur\u00adfaces a dif\u00adfer\u00adent class of fail\u00adure.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Jailbreak and adversarial probing<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Direct attempts to make the mod\u00adel pro\u00adduce con\u00adtent it should refuse: role\u00adplay and hypo\u00adthet\u00adi\u00adcal fram\u00adings, pay\u00adload smug\u00adgling, instruc\u00adtion over\u00adrides, and prompt injec\u00adtion. In a mul\u00adti\u00adlin\u00adgual set\u00adting, the same attack is authored fresh in each lan\u00adguage so that lan\u00adguage-spe\u00adcif\u00adic eva\u00adsions are caught.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Multilingual toxicity and harmful-content evaluation<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Mea\u00adsur\u00ading how often the mod\u00adel gen\u00ader\u00adates hate\u00adful, obscene, or oth\u00ader\u00adwise harm\u00adful text across lan\u00adguages, and how con\u00adsis\u00adtent\u00adly it refus\u00ades. This is where per-lan\u00adguage grad\u00ading mat\u00adters most, since a harm that is obvi\u00adous in Eng\u00adlish can be scored incon\u00adsis\u00adtent\u00adly in anoth\u00ader lan\u00adguage with\u00adout native review\u00aders. It con\u00adnects direct\u00adly to pro\u00adduc\u00adtion <a href=\"https:\/\/www.graveiensai.com\/content-moderation\">con\u00adtent mod\u00ader\u00ada\u00adtion<\/a>, because the harms test\u00aded here are the harms a live sys\u00adtem must catch.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Multilingual bias and fairness testing<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Check\u00ading whether the mod\u00adel treats peo\u00adple and groups dif\u00adfer\u00adent\u00adly depend\u00ading on the lan\u00adguage of the prompt or the group named in it: stereo\u00adtyp\u00ading, unequal refusal behav\u00adior, or skewed sen\u00adti\u00adment. Bias that is masked in Eng\u00adlish can sur\u00adface strong\u00adly in anoth\u00ader lan\u00adguage, so mul\u00adti\u00adlin\u00adgual bias and fair\u00adness test\u00ading is a sep\u00ada\u00adrate track rather than a byprod\u00aduct of tox\u00adi\u00adc\u00adi\u00adty work.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Multi-turn and code-switching attacks<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Esca\u00adlat\u00ading a con\u00adver\u00adsa\u00adtion over sev\u00ader\u00adal turns, or switch\u00ading lan\u00adguages with\u00adin a sin\u00adgle exchange, to erode safe\u00adguards that hold on the first, Eng\u00adlish, sin\u00adgle-turn prompt. Because vul\u00adner\u00ada\u00adbil\u00adi\u00adty ris\u00ades with con\u00adver\u00adsa\u00adtion length, sin\u00adgle-turn test\u00ading alone under\u00adstates real risk.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>How multilingual LLM red teaming works<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A well-run engage\u00adment moves through a repeat\u00adable sequence:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Scope and tax\u00adon\u00ado\u00admy.<\/strong> Agree on tar\u00adget lan\u00adguages, attack class\u00ades, harm cat\u00ade\u00adgories, and the grad\u00ading rubric before any probe is writ\u00adten, so results are com\u00adpa\u00adra\u00adble across lan\u00adguages.<\/li>\n\n\n\n<li><strong>Native-speak\u00ader probe design.<\/strong> Vet\u00adted native speak\u00aders author orig\u00adi\u00adnal adver\u00adsar\u00adi\u00adal and jail\u00adbreak prompts in each lan\u00adguage, cap\u00adtur\u00ading idiom, translit\u00ader\u00ada\u00adtion, and code-switch\u00ading.<\/li>\n\n\n\n<li><strong>Response col\u00adlec\u00adtion.<\/strong> The mod\u00adel answers every probe, in both sin\u00adgle-turn and mul\u00adti-turn form, with meta\u00adda\u00adta cap\u00adtured for trace\u00adabil\u00adi\u00adty.<\/li>\n\n\n\n<li><strong>Grad\u00ading.<\/strong> Native-speak\u00ader review\u00aders score each response for sever\u00adi\u00adty against the tax\u00adon\u00ado\u00admy and write a short ratio\u00adnale, so a fail\u00adure in one lan\u00adguage means the same as a fail\u00adure in anoth\u00ader.<\/li>\n\n\n\n<li><strong>Qual\u00adi\u00adty assur\u00adance.<\/strong> Labels pass a mul\u00adti-stage review with con\u00adsis\u00adten\u00adcy checks and gold-set audits before deliv\u00adery.<\/li>\n\n\n\n<li><strong>Report\u00ading and hand\u00adoff.<\/strong> The client receives a labeled dataset plus a per-lan\u00adguage cov\u00ader\u00adage and find\u00adings report, ready to dri\u00adve align\u00adment and guardrail work.<\/li>\n<\/ol>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>The Graveiens LENS Framework for language prioritization<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">No team can test every lan\u00adguage at full depth on day one, so the first real deci\u00adsion in mul\u00adti\u00adlin\u00adgual LLM red team\u00ading is which lan\u00adguages to test first. Most teams default to \u201cthe biggest mar\u00adkets,\u201d which qui\u00adet\u00adly ignores where mod\u00adels are most like\u00adly to fail. The LENS Frame\u00adwork scores each can\u00addi\u00addate lan\u00adguage or mar\u00adket on four dimen\u00adsions, each from 1 (low) to 5 (high). Add the scores for a pri\u00ador\u00adi\u00adty rat\u00ading from 4 to 20; test the high\u00adest scores first.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Dimen\u00adsion<\/strong><\/th><th><strong>What to eval\u00adu\u00adate<\/strong><\/th><th><strong>Score 1 to 5<\/strong><\/th><\/tr><\/thead><tbody><tr><td><strong>L: Lan\u00adguage reach<\/strong><\/td><td>How many users or how much rev\u00adenue depend on this lan\u00adguage in your prod\u00aduct.<\/td><td>1 = niche, 5 = core mar\u00adket<\/td><\/tr><tr><td><strong>E: Expo\u00adsure to harm<\/strong><\/td><td>Sever\u00adi\u00adty if the mod\u00adel fails here: reg\u00adu\u00adlat\u00aded domain, vul\u00adner\u00ada\u00adble users, safe\u00adty-crit\u00adi\u00adcal use.<\/td><td>1 = low stakes, 5 = high stakes<\/td><\/tr><tr><td><strong>N: Norms and lin\u00adguis\u00adtic com\u00adplex\u00adi\u00adty<\/strong><\/td><td>How much code-switch\u00ading, translit\u00ader\u00ada\u00adtion, dialect, and script mix\u00ading real users bring, all of which machine trans\u00adla\u00adtion miss\u00ades.<\/td><td>1 = sim\u00adple, 5 = high\u00adly mixed<\/td><\/tr><tr><td><strong>S: Safe\u00adty-data scarci\u00adty<\/strong><\/td><td>How low-resource the lan\u00adguage is, since thin\u00adner safe\u00adty train\u00ading means weak\u00ader guardrails.<\/td><td>1 = high-resource, 5 = low-resource<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The frame\u00adwork is delib\u00ader\u00adate\u00adly biased toward the lan\u00adguages the indus\u00adtry tends to skip. A high-traf\u00adfic lan\u00adguage with heavy code-switch\u00ading and thin safe\u00adty data, com\u00admon across South Asian and African mar\u00adkets, will score high\u00ader than a large but well-resourced Euro\u00adpean lan\u00adguage, which match\u00ades the research find\u00ading that risk con\u00adcen\u00adtrates in low-resource and mixed-script set\u00adtings. LENS decides the order of work; a full cov\u00ader\u00adage plan then tests each cho\u00adsen lan\u00adguage across lan\u00adguage tier, attack class, and turn depth.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Approaches compared: which method wins when<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">There is no sin\u00adgle best way to red team a mul\u00adti\u00adlin\u00adgual mod\u00adel. The real\u00adis\u00adtic ques\u00adtion is which method to use for which pur\u00adpose, and how to com\u00adbine them.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Approach<\/strong><\/th><th><strong>Best for<\/strong><\/th><th><strong>Strengths<\/strong><\/th><th><strong>Lim\u00adi\u00adta\u00adtions<\/strong><\/th><th><strong>Cost pro\u00adfile<\/strong><\/th><\/tr><\/thead><tbody><tr><td>Auto\u00admat\u00aded scan\u00adners<\/td><td>Fast, repeat\u00adable regres\u00adsion checks<\/td><td>Cheap, high vol\u00adume, runs on every build<\/td><td>Eng\u00adlish-cen\u00adtric, blind to cul\u00adtur\u00adal and code-switch\u00ading attacks, shal\u00adlow on nov\u00adel harms<\/td><td>Low per run<\/td><\/tr><tr><td>Machine-trans\u00adlat\u00aded human review<\/td><td>A rough first look at a new lan\u00adguage<\/td><td>Faster than author\u00ading from scratch<\/td><td>Miss\u00ades idiom, translit\u00ader\u00ada\u00adtion, and code-switch\u00ading, so it under-reports real risk<\/td><td>Low to mod\u00ader\u00adate<\/td><\/tr><tr><td>Eng\u00adlish-only human red team\u00ading<\/td><td>Prod\u00aducts that tru\u00adly ship in Eng\u00adlish only<\/td><td>Real human cre\u00adativ\u00adi\u00adty and sever\u00adi\u00adty judge\u00adment<\/td><td>No vis\u00adi\u00adbil\u00adi\u00adty into non-Eng\u00adlish fail\u00adure modes<\/td><td>Mod\u00ader\u00adate<\/td><\/tr><tr><td>Mul\u00adti\u00adlin\u00adgual native-speak\u00ader red team\u00ading<\/td><td>Mod\u00adels ship\u00adping in mul\u00adti\u00adple lan\u00adguages<\/td><td>Catch\u00ades the safe-in-Eng\u00adlish, bro\u00adken-else\u00adwhere gap; labels are reusable for align\u00adment<\/td><td>Needs a vet\u00adted native-speak\u00ader net\u00adwork and struc\u00adtured QA<\/td><td>High\u00ader, high\u00adest sig\u00adnal<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Auto\u00admat\u00aded scan\u00adners are usu\u00adal\u00adly stronger for catch\u00ading regres\u00adsions cheap\u00adly on every release, and machine trans\u00adla\u00adtion can be accept\u00adable for a quick sniff test. Native-speak\u00ader test\u00ading is prefer\u00adable when\u00adev\u00ader a real fail\u00adure in anoth\u00ader lan\u00adguage would harm users or breach an oblig\u00ada\u00adtion, which is most con\u00adsumer and enter\u00adprise deploy\u00adments. A hybrid approach tends to make sense: run scan\u00adners con\u00adtin\u00adu\u00adous\u00adly, then com\u00admis\u00adsion native-speak\u00ader pro\u00adgrams for the lan\u00adguages LENS ranks high\u00adest. The trade\u00adoff is straight\u00adfor\u00adward, cheap\u00ader meth\u00adods cost less per run but leave the high\u00adest-risk fail\u00adures unde\u00adtect\u00aded.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>The multilingual safety testing maturity model<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Teams rarely jump straight to full cov\u00ader\u00adage. This matu\u00adri\u00adty mod\u00adel helps you locate your cur\u00adrent stage and plan the next one.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Lev\u00adel<\/strong><\/th><th><strong>State<\/strong><\/th><th><strong>What it means<\/strong><\/th><\/tr><\/thead><tbody><tr><td>0<\/td><td>Eng\u00adlish-only<\/td><td>Safe\u00adty is mea\u00adsured in Eng\u00adlish and assumed to hold else\u00adwhere.<\/td><\/tr><tr><td>1<\/td><td>Trans\u00adlat\u00aded probes<\/td><td>Eng\u00adlish tests are machine-trans\u00adlat\u00aded; risk is under-report\u00aded.<\/td><\/tr><tr><td>2<\/td><td>Native probes, top lan\u00adguages<\/td><td>Orig\u00adi\u00adnal probes authored in the high\u00adest-pri\u00ador\u00adi\u00adty lan\u00adguages, sin\u00adgle-turn.<\/td><\/tr><tr><td>3<\/td><td>Mul\u00adti-tier, mul\u00adti-turn<\/td><td>Native probes across lan\u00adguage tiers, includ\u00ading mul\u00adti-turn and code-switch\u00ading, with grad\u00aded ratio\u00adnale.<\/td><\/tr><tr><td>4<\/td><td>Con\u00adtin\u00adu\u00adous and aligned<\/td><td>Per-lan\u00adguage test\u00ading runs on a sched\u00adule and feeds align\u00adment and guardrails as a closed loop.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Most orga\u00adni\u00adza\u00adtions ship\u00adping glob\u00adal\u00adly sit at Lev\u00adel 0 or 1 and believe they are fur\u00adther along. Mov\u00ading to Lev\u00adel 2 for even three or four high-pri\u00ador\u00adi\u00adty lan\u00adguages usu\u00adal\u00adly deliv\u00aders the largest sin\u00adgle jump in real safe\u00adty assur\u00adance.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>What multilingual red teaming costs<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Pric\u00ading depends on scope, so the hon\u00adest answer is that it is quot\u00aded per pro\u00adgram rather than sold at a fixed rate. What you can esti\u00admate in advance is the shape of the cost. The main dri\u00advers com\u00adbine like this:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Total cost is approx\u00adi\u00admate\u00adly: (num\u00adber of probes) x (num\u00adber of lan\u00adguages) x (per-probe author\u00ading and grad\u00ading effort) + pro\u00adgram man\u00adage\u00adment + QA over\u00adhead.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The fol\u00adlow\u00ading fig\u00adures are an illus\u00adtra\u00adtive mod\u00adel, not a quote, to show how scope moves the total. Sup\u00adpose a pro\u00adgram cov\u00aders 6 lan\u00adguages, 500 native probes per lan\u00adguage, each probe authored and grad\u00aded once, plus mul\u00adti-turn fol\u00adlow-ups on a sub\u00adset. The probe count alone is 3,000, before mul\u00adti-turn expan\u00adsion and QA. Dou\u00adbling the lan\u00adguage count rough\u00adly dou\u00adbles author\u00ading and grad\u00ading effort; adding mul\u00adti-turn depth increas\u00ades grad\u00ading effort per probe rather than probe count. Because the high\u00adest-risk lan\u00adguages are often low-resource, native-speak\u00ader sup\u00adply is the real con\u00adstraint on both cost and time\u00adline, which is anoth\u00ader rea\u00adson to pri\u00ador\u00adi\u00adtize with LENS rather than test\u00ading every\u00adthing shal\u00adlow\u00adly. For a scoped esti\u00admate on your own lan\u00adguages and harms, request a quote rather than rely\u00ading on a gener\u00adic fig\u00adure.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Common mistakes and how to avoid them<\/strong><\/h2>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Treat\u00ading machine-trans\u00adlat\u00aded Eng\u00adlish probes as mul\u00adti\u00adlin\u00adgual cov\u00ader\u00adage.<\/strong> It hap\u00adpens because trans\u00adla\u00adtion is fast and cheap. It mat\u00adters because it sys\u00adtem\u00adat\u00adi\u00adcal\u00adly under-reports risk. Pre\u00advent it by author\u00ading orig\u00adi\u00adnal probes in each lan\u00adguage.<\/li>\n\n\n\n<li><strong>Test\u00ading only sin\u00adgle-turn prompts.<\/strong> Teams do this because sin\u00adgle-turn is easy to auto\u00admate. But vul\u00adner\u00ada\u00adbil\u00adi\u00adty ris\u00ades over a con\u00adver\u00adsa\u00adtion, so sin\u00adgle-turn results over\u00adstate safe\u00adty. Add mul\u00adti-turn probes to your scope.<\/li>\n\n\n\n<li><strong>Ignor\u00ading low-resource lan\u00adguages.<\/strong> These are skipped because data and review\u00aders are scarce, yet they are where safe\u00adty train\u00ading is thinnest and jail\u00adbreaks suc\u00adceed most. Pri\u00ador\u00adi\u00adtize them explic\u00adit\u00adly.<\/li>\n\n\n\n<li><strong>Con\u00adfus\u00ading red team\u00ading with cer\u00adti\u00adfi\u00adca\u00adtion.<\/strong> Red team\u00ading pro\u00adduces evi\u00addence of fail\u00adure; it does not make a mod\u00adel safe by itself. Treat the find\u00adings as the input to align\u00adment work, not the fin\u00adish line.<\/li>\n\n\n\n<li><strong>Cap\u00adtur\u00ading labels with\u00adout ratio\u00adnale.<\/strong> A bare pass or fail is hard to reuse. Require sever\u00adi\u00adty plus a writ\u00adten rea\u00adson so labels feed both eval\u00adu\u00ada\u00adtion and align\u00adment train\u00ading.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">Also read: <a href=\"https:\/\/www.graveiensai.com\/blog\/content-moderation-services\">Con\u00adtent Mod\u00ader\u00ada\u00adtion Ser\u00advices: Types, Costs and How to Choose<\/a>, for how the harms you red team for map to live mod\u00ader\u00ada\u00adtion.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>A practical multilingual red-teaming checklist<\/strong><\/h2>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Define the harm tax\u00adon\u00ado\u00admy and grad\u00ading rubric you will use across all lan\u00adguages.<\/li>\n\n\n\n<li>Pri\u00ador\u00adi\u00adtize lan\u00adguages with the LENS Frame\u00adwork and pick your first cohort.<\/li>\n\n\n\n<li>Set attack class\u00ades to cov\u00ader: jail\u00adbreak, prompt injec\u00adtion, unsafe instruc\u00adtion, harm\u00adful con\u00adtent.<\/li>\n\n\n\n<li>Spec\u00adi\u00adfy turn depth: run both sin\u00adgle-turn and mul\u00adti-turn, plus code-switch\u00ading cas\u00ades.<\/li>\n\n\n\n<li>Recruit and vet native-speak\u00ader probe writ\u00aders and review\u00aders per lan\u00adguage.<\/li>\n\n\n\n<li>Author orig\u00adi\u00adnal probes; do not trans\u00adlate an Eng\u00adlish set.<\/li>\n\n\n\n<li>Col\u00adlect respons\u00ades with meta\u00adda\u00adta for trace\u00adabil\u00adi\u00adty.<\/li>\n\n\n\n<li>Grade every response for sever\u00adi\u00adty with a writ\u00adten ratio\u00adnale.<\/li>\n\n\n\n<li>Run mul\u00adti-stage QA with con\u00adsis\u00adten\u00adcy checks and gold-set audits.<\/li>\n\n\n\n<li>Deliv\u00ader a per-lan\u00adguage find\u00adings report and route fail\u00adures into align\u00adment and guardrails.<\/li>\n\n\n\n<li>Re-test on a sched\u00adule and after every major mod\u00adel update.<\/li>\n<\/ol>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Worked examples<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The fol\u00adlow\u00ading are illus\u00adtra\u00adtive sce\u00adnar\u00adios, not spe\u00adcif\u00adic cus\u00adtomer results.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Exam\u00adple 1: A con\u00adsumer chat\u00adbot expand\u00ading into India.<\/strong> Before launch, safe\u00adty was mea\u00adsured in Eng\u00adlish and passed. Prob\u00adlem: the prod\u00aduct would serve Hin\u00addi, Ben\u00adgali, and Tamil users who rou\u00adtine\u00adly mix Eng\u00adlish and local scripts. Deci\u00adsion: run native-speak\u00ader red team\u00ading on those three lan\u00adguages, includ\u00ading translit\u00ader\u00adat\u00aded and code-switched prompts. Imple\u00admen\u00adta\u00adtion: orig\u00adi\u00adnal jail\u00adbreak and harm\u00adful-con\u00adtent probes per lan\u00adguage, grad\u00aded with ratio\u00adnale. Expect\u00aded out\u00adcome: fail\u00adure modes invis\u00adi\u00adble in Eng\u00adlish are sur\u00adfaced and fixed before launch, and the grad\u00aded data seeds Indic-lan\u00adguage align\u00adment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Exam\u00adple 2: An enter\u00adprise mod\u00adel with a com\u00adpli\u00adance oblig\u00ada\u00adtion.<\/strong> A mod\u00adel deployed across sev\u00ader\u00adal Euro\u00adpean and Mid\u00addle East\u00adern mar\u00adkets faces adver\u00adsar\u00adi\u00adal-test\u00ading expec\u00adta\u00adtions for high\u00ader-risk sys\u00adtems. Prob\u00adlem: an Eng\u00adlish-only eval\u00adu\u00ada\u00adtion will not sat\u00adis\u00adfy a reg\u00adu\u00adla\u00adtor ask\u00ading about the lan\u00adguages actu\u00adal\u00adly served. Deci\u00adsion: com\u00admis\u00adsion a mul\u00adti\u00adlin\u00adgual pro\u00adgram cov\u00ader\u00ading the deployed lan\u00adguages across attack class\u00ades and turn depth, deliv\u00adered as a defen\u00adsi\u00adble per-lan\u00adguage cov\u00ader\u00adage report. Expect\u00aded out\u00adcome: doc\u00adu\u00adment\u00aded evi\u00addence of where the mod\u00adel was test\u00aded and how it per\u00adformed, ready for inter\u00adnal gov\u00ader\u00adnance and exter\u00adnal review.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>How regulation is raising the bar<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Adver\u00adsar\u00adi\u00adal test\u00ading is mov\u00ading from best prac\u00adtice toward expec\u00adta\u00adtion. Under the EU AI Act, providers of gen\u00ader\u00adal-pur\u00adpose AI mod\u00adels with sys\u00adtemic risk are expect\u00aded to per\u00adform adver\u00adsar\u00adi\u00adal test\u00ading, com\u00admon\u00adly under\u00adstood as mod\u00adel eval\u00adu\u00ada\u00adtion and red team\u00ading, to iden\u00adti\u00adfy and mit\u00adi\u00adgate sys\u00adtemic risks. The US NIST AI Risk Man\u00adage\u00adment Frame\u00adwork and its gen\u00ader\u00ada\u00adtive AI pro\u00adfile treat struc\u00adtured red team\u00ading as a core prac\u00adtice, and the OWASP Top 10 for LLM Appli\u00adca\u00adtions cat\u00ada\u00adlogs the vul\u00adner\u00ada\u00adbil\u00adi\u00adty class\u00ades, such as prompt injec\u00adtion, that red teams probe. None of these frame\u00adworks says Eng\u00adlish is enough. For any provider serv\u00ading mul\u00adti\u00adple lan\u00adguages, defen\u00adsi\u00adble test\u00ading means test\u00ading in the lan\u00adguages the mod\u00adel actu\u00adal\u00adly serves, which is the premise of mul\u00adti\u00adlin\u00adgual red team\u00ading. Con\u00adfirm the cur\u00adrent text of any reg\u00adu\u00adla\u00adtion before rely\u00ading on it for a com\u00adpli\u00adance deci\u00adsion, since these rules are still being imple\u00adment\u00aded.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Frequently asked questions<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>What is mul\u00adti\u00adlin\u00adgual LLM red team\u00ading?<\/strong><br>It is adver\u00adsar\u00adi\u00adal safe\u00adty test\u00ading of a large lan\u00adguage mod\u00adel across every lan\u00adguage it serves, per\u00adformed by native speak\u00aders who write orig\u00adi\u00adnal attack prompts and grade the mod\u00adel\u2019s respons\u00ades against a harm tax\u00adon\u00ado\u00admy. The goal is to find where safe\u00adguards fail out\u00adside Eng\u00adlish, then feed that evi\u00addence into eval\u00adu\u00ada\u00adtion and align\u00adment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How is red team\u00ading dif\u00adfer\u00adent from LLM eval\u00adu\u00ada\u00adtion?<\/strong><br>Red team\u00ading is adver\u00adsar\u00adi\u00adal: testers active\u00adly try to make the mod\u00adel fail. Eval\u00adu\u00ada\u00adtion is broad\u00ader and often mea\u00adsures gen\u00ader\u00adal qual\u00adi\u00adty or capa\u00adbil\u00adi\u00adty. Red-team find\u00adings are a spe\u00adcial\u00adized, safe\u00adty-focused input to a wider eval\u00adu\u00ada\u00adtion pro\u00adgram, and the two work togeth\u00ader.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Why do lan\u00adguage mod\u00adels fail more in non-Eng\u00adlish lan\u00adguages?<\/strong><br>Because safe\u00adty align\u00adment is trained most\u00adly on Eng\u00adlish data. Low\u00ader-resource lan\u00adguages get less safe\u00adty train\u00ading, so their guardrails are weak\u00ader. Stud\u00adies have mea\u00adsured rough\u00adly three times the harm\u00adful-out\u00adput rate in low-resource lan\u00adguages com\u00adpared with high-resource ones.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Can we just machine-trans\u00adlate our Eng\u00adlish red-team set?<\/strong><br>It is bet\u00adter than noth\u00ading but not suf\u00adfi\u00adcient. Trans\u00adla\u00adtion miss\u00ades idioms, translit\u00ader\u00ada\u00adtion, and code-switch\u00ading, the pat\u00adterns real users and attack\u00aders actu\u00adal\u00adly use, so it under-reports true risk. Native-authored probes are need\u00aded for reli\u00adable cov\u00ader\u00adage.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>What harms does mul\u00adti\u00adlin\u00adgual red team\u00ading cov\u00ader?<\/strong><br>Typ\u00adi\u00adcal\u00adly jail\u00adbreaks and adver\u00adsar\u00adi\u00adal attacks, mul\u00adti\u00adlin\u00adgual tox\u00adi\u00adc\u00adi\u00adty and harm\u00adful-con\u00adtent eval\u00adu\u00ada\u00adtion, mul\u00adti\u00adlin\u00adgual bias and fair\u00adness test\u00ading, and prompt injec\u00adtion, test\u00aded across sin\u00adgle-turn and mul\u00adti-turn con\u00adver\u00adsa\u00adtions and, where rel\u00ade\u00advant, code-switch\u00ading.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How much does it cost?<\/strong><br>There is no fixed price; cost scales with the num\u00adber of probes, the num\u00adber of lan\u00adguages, grad\u00ading depth, and QA. Because the high\u00adest-risk lan\u00adguages are often low-resource, native-speak\u00ader avail\u00adabil\u00adi\u00adty is the main con\u00adstraint. Pro\u00adgrams are scoped and quot\u00aded per engage\u00adment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>How do we choose which lan\u00adguages to test first?<\/strong><br>Pri\u00ador\u00adi\u00adtize by expo\u00adsure and risk, not just mar\u00adket size. The LENS Frame\u00adwork scores each lan\u00adguage on reach, harm expo\u00adsure, lin\u00adguis\u00adtic com\u00adplex\u00adi\u00adty, and safe\u00adty-data scarci\u00adty, which tends to sur\u00adface high-traf\u00adfic, low-resource, code-switch\u00ading lan\u00adguages that are oth\u00ader\u00adwise skipped.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Does red team\u00ading make a mod\u00adel safe?<\/strong><br>No. It pro\u00adduces evi\u00addence of where the mod\u00adel fails. Mak\u00ading the mod\u00adel safer is a sep\u00ada\u00adrate step, using the grad\u00aded find\u00adings to dri\u00adve align\u00adment train\u00ading and run\u00adtime guardrails, fol\u00adlowed by re-test\u00ading.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Who should run mul\u00adti\u00adlin\u00adgual red team\u00ading?<\/strong><br>AI lab safe\u00adty teams ship\u00adping into non-Eng\u00adlish mar\u00adkets, third-par\u00adty eval\u00adu\u00ada\u00adtors and AI Safe\u00adty Insti\u00adtutes, and enter\u00adpris\u00ades deploy\u00ading mod\u00adels in mul\u00adti\u00adlin\u00adgual regions. It requires a vet\u00adted native-speak\u00ader net\u00adwork and struc\u00adtured qual\u00adi\u00adty assur\u00adance, which is why many teams part\u00adner with a spe\u00adcial\u00adized <a href=\"https:\/\/www.graveiensai.com\/generative-ai\">gen\u00ader\u00ada\u00adtive AI data<\/a> provider.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>About the authors<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This guide was writ\u00adten by the Graveiens AI edi\u00adto\u00adr\u00adi\u00adal team, led by Jiten\u00addra Choubay, Founder and CEO of Graveiens AI. Graveiens AI is a human-in-the-loop AI data com\u00adpa\u00adny that sup\u00adplies con\u00adsent-backed adver\u00adsar\u00adi\u00adal data and harm rat\u00adings for LLM safe\u00adty pro\u00adgrams, with native-speak\u00ader cov\u00ader\u00adage across 25 or more lan\u00adguages includ\u00ading deep Indic sup\u00adport, backed by a four-stage qual\u00adi\u00adty assur\u00adance work\u00adflow and ISO 9001:2017 qual\u00adi\u00adty man\u00adage\u00adment. [Review\u00ader name, title, and rel\u00ade\u00advant AI safe\u00adty or NLP cre\u00adden\u00adtials to be con\u00adfirmed by Om before pub\u00adli\u00adca\u00adtion.] Learn more about the team and method\u00adol\u00ado\u00adgy on the <a href=\"https:\/\/www.graveiensai.com\/\">Graveiens AI<\/a> site.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Conclusion<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Mul\u00adti\u00adlin\u00adgual LLM red team\u00ading is how you find out whether a mod\u00adel is actu\u00adal\u00adly safe for the peo\u00adple who use it, rather than safe only for the peo\u00adple who test\u00aded it. The evi\u00addence is con\u00adsis\u00adtent across inde\u00adpen\u00addent stud\u00adies: safe\u00adguards that hold in Eng\u00adlish degrade sharply in low\u00ader-resource lan\u00adguages, get worse over mul\u00adti-turn con\u00adver\u00adsa\u00adtions, and are missed entire\u00adly by machine-trans\u00adlat\u00aded tests. The path for\u00adward is prac\u00adti\u00adcal. Pri\u00ador\u00adi\u00adtize lan\u00adguages with a clear frame\u00adwork such as LENS, run mul\u00adti\u00adlin\u00adgual tox\u00adi\u00adc\u00adi\u00adty and harm\u00adful-con\u00adtent eval\u00adu\u00ada\u00adtion along\u00adside mul\u00adti\u00adlin\u00adgual bias and fair\u00adness test\u00ading, cov\u00ader attack class\u00ades and turn depth with native speak\u00aders, grade with ratio\u00adnale, and route the find\u00adings into align\u00adment and guardrails, then re-test. If you need native-speak\u00ader adver\u00adsar\u00adi\u00adal data and harm eval\u00adu\u00ada\u00adtion across the lan\u00adguages your mod\u00adel serves, Graveiens AI can scope a <a href=\"https:\/\/www.graveiensai.com\/ai-red-teaming-services\">mul\u00adti\u00adlin\u00adgual red-team\u00ading pilot<\/a> you can judge on your own mod\u00adel.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Sources<\/strong><\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Deng et al., \u201cMul\u00adti\u00adlin\u00adgual Jail\u00adbreak Chal\u00adlenges in Large Lan\u00adguage Mod\u00adels,\u201d ICLR 2024. <a href=\"https:\/\/arxiv.org\/html\/2310.06474v3\" target=\"_blank\" rel=\"noopener\">arxiv.org\/html\/2310.06474v3<\/a><\/li>\n\n\n\n<li>Jain et al., \u201cPoly\u00adglo\u00adTox\u00adi\u00adc\u00adi\u00adtyPrompts: Mul\u00adti\u00adlin\u00adgual Eval\u00adu\u00ada\u00adtion of Neur\u00adal Tox\u00adic Degen\u00ader\u00ada\u00adtion in Large Lan\u00adguage Mod\u00adels,\u201d 2024. <a href=\"https:\/\/arxiv.org\/html\/2405.09373\" target=\"_blank\" rel=\"noopener\">arxiv.org\/html\/2405.09373<\/a><\/li>\n\n\n\n<li>Sing\u00adha\u00adnia et al. (Ama\u00adzon Sci\u00adence), \u201cMul\u00adti-lin\u00adgual Mul\u00adti-turn Auto\u00admat\u00aded Red Team\u00ading for LLMs,\u201d 2025. <a href=\"https:\/\/arxiv.org\/html\/2504.03174v1\" target=\"_blank\" rel=\"noopener\">arxiv.org\/html\/2504.03174v1<\/a><\/li>\n\n\n\n<li>Johns Hop\u00adkins Uni\u00adver\u00adsi\u00adty, \u201cJail\u00adbreaks Threat\u00aden Low-Resource Lan\u00adguages,\u201d 2024. <a href=\"https:\/\/engineering.jhu.edu\/magazine-archive\/2024\/12\/jailbreaks-threaten-low-resource-languages\" target=\"_blank\" rel=\"noopener\">engineering.jhu.edu<\/a><\/li>\n\n\n\n<li>Euro\u00adpean Union, \u201cEU Arti\u00adfi\u00adcial Intel\u00adli\u00adgence Act\u201d (oblig\u00ada\u00adtions for gen\u00ader\u00adal-pur\u00adpose AI mod\u00adels with sys\u00adtemic risk). <a href=\"https:\/\/artificialintelligenceact.eu\/\" target=\"_blank\" rel=\"noopener\">artificialintelligenceact.eu<\/a><\/li>\n\n\n\n<li>NIST, \u201cAI Risk Man\u00adage\u00adment Frame\u00adwork\u201d and Gen\u00ader\u00ada\u00adtive AI Pro\u00adfile. <a href=\"https:\/\/www.nist.gov\/itl\/ai-risk-management-framework\" target=\"_blank\" rel=\"noopener\">nist.gov<\/a><\/li>\n\n\n\n<li>OWASP, \u201cTop 10 for Large Lan\u00adguage Mod\u00adel Appli\u00adca\u00adtions.\u201d <a href=\"https:\/\/genai.owasp.org\/\" target=\"_blank\" rel=\"noopener\">genai.owasp.org<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Mul\u00adti\u00adlin\u00adgual LLM red team\u00ading is the prac\u00adtice of adver\u00adsar\u00adi\u00adal\u00adly test\u00ading a large lan\u00adguage mod\u00adel for unsafe behav\u00adior in every lan\u00adguage it serves, not only in Eng\u00adlish, using native\u2026<\/p>\n","protected":false},"author":1,"featured_media":163,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"wp_typography_post_enhancements_disabled":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-162","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog"],"_links":{"self":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts\/162","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/comments?post=162"}],"version-history":[{"count":1,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts\/162\/revisions"}],"predecessor-version":[{"id":164,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts\/162\/revisions\/164"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/media\/163"}],"wp:attachment":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/media?parent=162"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/categories?post=162"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/tags?post=162"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}