{"id":105,"date":"2026-08-17T07:51:09","date_gmt":"2026-08-17T07:51:09","guid":{"rendered":"https:\/\/www.graveiensai.com\/blog\/?p=105"},"modified":"2026-08-17T07:51:09","modified_gmt":"2026-08-17T07:51:09","slug":"what-is-instruction-tuning","status":"publish","type":"post","link":"https:\/\/www.graveiensai.com\/blog\/what-is-instruction-tuning\/","title":{"rendered":"What Is Instruction Tuning? A Complete 2026 Guide"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Instruc\u00adtion tun\u00ading is a form of super\u00advised fine-tun\u00ading that trains a pre\u00adtrained lan\u00adguage mod\u00adel on instruc\u00adtion-and-response exam\u00adples so it learns to fol\u00adlow nat\u00adur\u00adal-lan\u00adguage com\u00admands. It teach\u00ades the mod\u00adel how to respond to instruc\u00adtions, rather than sim\u00adply pre\u00addict\u00ading the next token from gen\u00ader\u00adal web or text data.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In oth\u00ader words, instruc\u00adtion tun\u00ading is the step that turns a raw foun\u00adda\u00adtion mod\u00adel, which is a strong pat\u00adtern pre\u00addic\u00adtor, into a help\u00adful assis\u00adtant that does what you ask. Base mod\u00adels like the pre\u00adtrained ver\u00adsions of GPT, Lla\u00adma, or Gem\u00adi\u00adni under\u00adstand lan\u00adguage deeply but do not nat\u00adu\u00adral\u00adly fol\u00adlow instruc\u00adtions. Instruc\u00adtion tun\u00ading fix\u00ades that by train\u00ading the mod\u00adel on many exam\u00adples such as \u201cSum\u00admarise this para\u00adgraph\u201d paired with a good sum\u00adma\u00adry. After this step, the mod\u00adel gen\u00ader\u00adalis\u00ades to new instruc\u00adtions it has nev\u00ader seen.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Instruction tuning at a glance<\/strong><\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Ques\u00adtion<\/strong><\/th><th><strong>Short answer<\/strong><\/th><\/tr><\/thead><tbody><tr><td><strong>What is instruc\u00adtion tun\u00ading?<\/strong><\/td><td>A form of super\u00advised fine-tun\u00ading on instruc\u00adtion-and-response pairs so a mod\u00adel fol\u00adlows com\u00admands.<\/td><\/tr><tr><td><strong>Is it the same as super\u00advised fine-tun\u00ading?<\/strong><\/td><td>Instruc\u00adtion tun\u00ading is super\u00advised fine-tun\u00ading (SFT) done with instruc\u00adtion-for\u00admat\u00adted data.<\/td><\/tr><tr><td><strong>Why does it mat\u00adter?<\/strong><\/td><td>It turns a raw text pre\u00addic\u00adtor into a usable, instruc\u00adtion-fol\u00adlow\u00ading assis\u00adtant.<\/td><\/tr><tr><td><strong>Where does it sit?<\/strong><\/td><td>Between pre\u00adtrain\u00ading and pref\u00ader\u00adence opti\u00admiza\u00adtion such as RLHF.<\/td><\/tr><tr><td><strong>What does it need?<\/strong><\/td><td>A high-qual\u00adi\u00adty, diverse dataset of labelled instruc\u00adtion-response exam\u00adples.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>What is instruction tuning?<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To answer what is instruc\u00adtion tun\u00ading pre\u00adcise\u00adly: it is super\u00advised train\u00ading in which a pre\u00adtrained lan\u00adguage mod\u00adel is fur\u00adther trained on a curat\u00aded set of tasks, each described in nat\u00adur\u00adal lan\u00adguage, so it learns the gen\u00ader\u00adal skill of instruc\u00adtion-fol\u00adlow\u00ading rather than a sin\u00adgle nar\u00adrow task. The train\u00ading data is a col\u00adlec\u00adtion of instruc\u00adtion-response pairs, often writ\u00adten as an instruc\u00adtion, an option\u00adal input, and a desired out\u00adput.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The rea\u00adson this works is gen\u00ader\u00adal\u00adi\u00adsa\u00adtion. When a mod\u00adel sees many var\u00adied instruc\u00adtions dur\u00ading train\u00ading, from sum\u00adma\u00adriz\u00ading to clas\u00adsi\u00adfy\u00ading to writ\u00ading code, it does not just mem\u00ado\u00adrise those tasks; it learns the under\u00adly\u00ading behav\u00adior of read\u00ading an instruc\u00adtion and pro\u00adduc\u00ading a suit\u00adable response. This is close\u00adly relat\u00aded to zero-shot and few-shot learn\u00ading: a well instruc\u00adtion-tuned mod\u00adel can han\u00addle prompts it nev\u00ader saw in train\u00ading.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instruc\u00adtion tun\u00ading is a core ser\u00advice in the wider prac\u00adtice of <a href=\"https:\/\/www.graveiensai.com\/llm-fine\">LLM fine-tun\u00ading<\/a> and mod\u00adern <a href=\"https:\/\/www.graveiensai.com\/generative-ai\">gen\u00ader\u00ada\u00adtive AI<\/a>. It is the bridge between a mod\u00adel that can pre\u00addict text and a mod\u00adel that can be told what to do.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Also read: <\/strong>If you are new to the under\u00adly\u00ading tech\u00adnol\u00ado\u00adgy, our explain\u00ader on <a href=\"https:\/\/www.graveiensai.com\/blog\/what-is-llm\">what an LLM is<\/a> cov\u00aders how large lan\u00adguage mod\u00adels work before they are tuned.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>How instruction tuning works<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">At a mechan\u00adi\u00adcal lev\u00adel, instruc\u00adtion tun\u00ading is stan\u00addard super\u00advised fine-tun\u00ading applied to a spe\u00adcif\u00adic kind of data. The work\u00adflow has four stages.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>1. <\/strong><strong>Start with a pre\u00adtrained mod\u00adel <\/strong>that already under\u00adstands lan\u00adguage from large-scale pre\u00adtrain\u00ading.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>2. <\/strong><strong>Assem\u00adble an instruc\u00adtion dataset <\/strong>of exam\u00adples, each pair\u00ading a nat\u00adur\u00adal-lan\u00adguage instruc\u00adtion (and any input) with a high-qual\u00adi\u00adty tar\u00adget response.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>3. <\/strong><strong>Fine-tune <\/strong>the mod\u00adel on these pairs so its out\u00adput match\u00ades the tar\u00adget respons\u00ades. To save com\u00adpute, teams often use para\u00adme\u00adter-effi\u00adcient meth\u00adods such as LoRA and oth\u00ader PEFT tech\u00adniques instead of updat\u00ading every weight.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>4. <\/strong><strong>Eval\u00adu\u00adate <\/strong>the tuned mod\u00adel on held-out instruc\u00adtions to con\u00adfirm it gen\u00ader\u00adalis\u00ades.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The qual\u00adi\u00adty of the instruc\u00adtion dataset is one of the most impor\u00adtant fac\u00adtors deter\u00admin\u00ading how well the tuned mod\u00adel per\u00adforms. The mod\u00adel learns to imi\u00adtate the respons\u00ades it is shown, so if the exam\u00adples are accu\u00adrate, diverse, and well writ\u00adten, the mod\u00adel becomes help\u00adful. This is why dis\u00adci\u00adplined <a href=\"https:\/\/www.graveiensai.com\/data-annotation\">data anno\u00adta\u00adtion<\/a> and expert review sit at the heart of every suc\u00adcess\u00adful instruc\u00adtion-tun\u00ading project.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>An instruction tuning example<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A sin\u00adgle train\u00ading exam\u00adple makes the idea con\u00adcrete. Each exam\u00adple usu\u00adal\u00adly has three parts: an instruc\u00adtion, an option\u00adal input, and the ide\u00adal response the mod\u00adel should learn to pro\u00adduce.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Instruc\u00adtion: <\/strong>\u201cSum\u00admarise the fol\u00adlow\u00ading cus\u00adtomer com\u00adplaint in one sen\u00adtence.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Input: <\/strong>\u201cI ordered a lap\u00adtop stand on the 3rd, paid for express deliv\u00adery, and it still had not arrived by the 12th. Sup\u00adport did not reply to two emails, and I want a refund on the ship\u00adping at least.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Ide\u00adal response: <\/strong>\u201cThe cus\u00adtomer wants a ship\u00adping refund after a paid express order arrived nine days late with no sup\u00adport reply.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Dur\u00ading train\u00ading, the mod\u00adel sees thou\u00adsands or even mil\u00adlions of exam\u00adples like this, across many task types, and learns the gen\u00ader\u00adal behav\u00adiour of read\u00ading an instruc\u00adtion and return\u00ading the kind of answer the exam\u00adples demon\u00adstrate. Every response in the dataset must be cor\u00adrect and con\u00adsis\u00adtent\u00adly for\u00admat\u00adted. Pro\u00adduc\u00ading these human demon\u00adstra\u00adtions at scale, writ\u00adten and reviewed by skilled peo\u00adple, is exact\u00adly what our <a href=\"https:\/\/www.graveiensai.com\/workforce\">expert work\u00adforce<\/a> does.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Instruction tuning vs supervised fine-tuning<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Super\u00advised fine-tun\u00ading (SFT) is the gen\u00ader\u00adal method of train\u00ading a mod\u00adel on labelled input-out\u00adput pairs. Instruc\u00adtion tun\u00ading is super\u00advised fine-tun\u00ading where those pairs are for\u00admat\u00adted as instruc\u00adtions across many dif\u00adfer\u00adent tasks.<\/strong> In every\u00adday use, the two terms are often used inter\u00adchange\u00adably, and that is usu\u00adal\u00adly fine.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><\/th><th><strong>Super\u00advised fine-tun\u00ading (SFT)<\/strong><\/th><th><strong>Instruc\u00adtion tun\u00ading<\/strong><\/th><\/tr><\/thead><tbody><tr><td><strong>What it is<\/strong><\/td><td>The method: train\u00ading on labelled pairs<\/td><td>A pur\u00adpose: SFT using instruc\u00adtion data<\/td><\/tr><tr><td><strong>Data<\/strong><\/td><td>Any labelled exam\u00adples for a task<\/td><td>Diverse instruc\u00adtion-and-response pairs<\/td><\/tr><tr><td><strong>Goal<\/strong><\/td><td>Improve per\u00adfor\u00admance on tar\u00adget tasks<\/td><td>Teach gen\u00ader\u00adal instruc\u00adtion-fol\u00adlow\u00ading<\/td><\/tr><tr><td><strong>Scope<\/strong><\/td><td>Can be one nar\u00adrow task<\/td><td>Delib\u00ader\u00adate\u00adly many tasks<\/td><\/tr><tr><td><strong>Rela\u00adtion\u00adship<\/strong><\/td><td>The broad\u00ader cat\u00ade\u00adgo\u00adry<\/td><td>A spe\u00adcif\u00adic appli\u00adca\u00adtion of SFT<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Put sim\u00adply, all instruc\u00adtion tun\u00ading is SFT, but not all SFT is instruc\u00adtion tun\u00ading. You use instruc\u00adtion tun\u00ading when you want the mod\u00adel to fol\u00adlow open-end\u00aded instruc\u00adtions in gen\u00ader\u00adal. Because the mechan\u00adics are the same, teams that deliv\u00ader super\u00advised fine-tun\u00ading data can deliv\u00ader instruc\u00adtion-tun\u00ading data with the same pipeline, which is how our <a href=\"https:\/\/www.graveiensai.com\/llm-fine\">LLM fine-tun\u00ading<\/a> pro\u00adgrammes are struc\u00adtured.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Instruction tuning vs fine-tuning vs prompt engineering<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">It helps to place instruc\u00adtion tun\u00ading along\u00adside the oth\u00ader ways of shap\u00ading a mod\u00adel. Each one changes some\u00adthing dif\u00adfer\u00adent, at a dif\u00adfer\u00adent point in a mod\u00adel life.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Method<\/strong><\/th><th><strong>When it hap\u00adpens<\/strong><\/th><th><strong>What changes<\/strong><\/th><th><strong>Main pur\u00adpose<\/strong><\/th><\/tr><\/thead><tbody><tr><td><strong>Pre\u00adtrain\u00ading<\/strong><\/td><td>Before deploy\u00adment<\/td><td>Mod\u00adel weights<\/td><td>Learn lan\u00adguage and world pat\u00adterns<\/td><\/tr><tr><td><strong>Instruc\u00adtion tun\u00ading<\/strong><\/td><td>Train\u00ading (post-train\u00ading)<\/td><td>Mod\u00adel weights<\/td><td>Teach the mod\u00adel to fol\u00adlow instruc\u00adtions<\/td><\/tr><tr><td><strong>Fine-tun\u00ading<\/strong><\/td><td>Train\u00ading<\/td><td>Mod\u00adel weights<\/td><td>Adapt to spe\u00adcif\u00adic tasks or domains<\/td><\/tr><tr><td><strong>Prompt engi\u00adneer\u00ading<\/strong><\/td><td>Infer\u00adence (run time)<\/td><td>Prompt and con\u00adtext<\/td><td>Guide the out\u00adput of a trained mod\u00adel<\/td><\/tr><tr><td><strong>RLHF \/ pref\u00ader\u00adence opti\u00admiza\u00adtion<\/strong><\/td><td>Post-train\u00ading<\/td><td>Behav\u00adiour and weights<\/td><td>Align respons\u00ades with human pref\u00ader\u00adences<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The key dis\u00adtinc\u00adtion is when each method acts. Instruc\u00adtion tun\u00ading and fine-tun\u00ading change the mod\u00adel weights dur\u00ading train\u00ading, while prompt engi\u00adneer\u00ading changes only the input at run time. That is why the two are com\u00adple\u00admen\u00adtary: a well instruc\u00adtion-tuned mod\u00adel still ben\u00ade\u00adfits from good prompts. For the run-time side, see our guide to <a href=\"https:\/\/www.graveiensai.com\/blog\/what-is-prompt-engineering\">what prompt engi\u00adneer\u00ading is<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Where instruction tuning sits: pretraining, SFT, RLHF<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>1. <\/strong><strong>Pre\u00adtrain\u00ading: <\/strong>the mod\u00adel learns lan\u00adguage and world knowl\u00adedge from a mas\u00adsive text cor\u00adpus. It can pre\u00addict text but does not reli\u00adably fol\u00adlow instruc\u00adtions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>2. <\/strong><strong>Instruc\u00adtion tun\u00ading (SFT): <\/strong>the mod\u00adel is fine-tuned on instruc\u00adtion-response pairs and learns to fol\u00adlow com\u00admands. It is now gen\u00aduine\u00adly use\u00adful.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>3. <\/strong><strong>RLHF and pref\u00ader\u00adence opti\u00admiza\u00adtion: <\/strong>the mod\u00adel is refined using human pref\u00ader\u00adence data so its answers are more help\u00adful, safe, and aligned with what peo\u00adple want.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An impor\u00adtant point often missed: instruc\u00adtion tun\u00ading alone gets you sur\u00adpris\u00ading\u00adly far. Many instruc\u00adtion-fol\u00adlow\u00ading mod\u00adels use super\u00advised instruc\u00adtion tun\u00ading as an impor\u00adtant post-train\u00ading stage, with addi\u00adtion\u00adal pref\u00ader\u00adence opti\u00admiza\u00adtion or rein\u00adforce\u00adment-learn\u00ading-based meth\u00adods such as RLHF applied in some train\u00ading pipelines. Land\u00admark sys\u00adtems show the pat\u00adtern: Google FLAN demon\u00adstrat\u00aded that instruc\u00adtion tun\u00ading across many tasks sharply improves zero-shot gen\u00ader\u00adal\u00adi\u00adsa\u00adtion, while Ope\u00adnAI Instruct\u00adG\u00adPT com\u00adbined super\u00advised instruc\u00adtion tun\u00ading with RLHF and became the direct ances\u00adtor of Chat\u00adG\u00adPT. Eval\u00adu\u00adat\u00ading each stage requires struc\u00adtured <a href=\"https:\/\/www.graveiensai.com\/llm-evaluation\">LLM eval\u00adu\u00ada\u00adtion<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Why instruction tuning matters<\/strong><\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Usabil\u00adi\u00adty: <\/strong>it turns a raw text pre\u00addic\u00adtor into an assis\u00adtant that fol\u00adlows instruc\u00adtions.<\/li>\n\n\n\n<li><strong>Gen\u00ader\u00adal\u00adi\u00adsa\u00adtion: <\/strong>a mod\u00adel tuned on diverse tasks han\u00addles new, unseen instruc\u00adtions.<\/li>\n\n\n\n<li><strong>Con\u00adtrol: <\/strong>it lets teams shape tone, for\u00admat, safe\u00adty behav\u00adiour, and domain focus.<\/li>\n\n\n\n<li><strong>Effi\u00adcien\u00adcy: <\/strong>far cheap\u00ader than pre\u00adtrain\u00ading from scratch, espe\u00adcial\u00adly with PEFT meth\u00adods like LoRA that tune only a small frac\u00adtion of the weights.<\/li>\n\n\n\n<li><strong>Cus\u00adtomi\u00adsa\u00adtion: <\/strong>organ\u00adi\u00adsa\u00adtions can instruc\u00adtion-tune an open mod\u00adel on their own tasks to build a pri\u00advate assis\u00adtant.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">For many teams, instruc\u00adtion tun\u00ading can be a cost-effec\u00adtive way to adapt a capa\u00adble base mod\u00adel, par\u00adtic\u00adu\u00adlar\u00adly when com\u00adbined with retrieval, eval\u00adu\u00ada\u00adtion, and appro\u00adpri\u00adate pref\u00ader\u00adence-opti\u00admiza\u00adtion meth\u00adods. It is also a com\u00admon foun\u00adda\u00adtion for domain assis\u00adtants in health\u00adcare, finance, and law, where pre\u00adcise instruc\u00adtion-fol\u00adlow\u00ading mat\u00adters most.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Instruction tuning benefits and limitations<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The ben\u00ade\u00adfits above are real, but a bal\u00adanced view mat\u00adters. Instruc\u00adtion tun\u00ading has clear lim\u00adi\u00adta\u00adtions that every team should plan for:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>It does not auto\u00admat\u00adi\u00adcal\u00adly add new fac\u00adtu\u00adal knowl\u00adedge. <\/strong>Tun\u00ading teach\u00ades behav\u00adiour, not facts; new knowl\u00adedge usu\u00adal\u00adly comes from pre\u00adtrain\u00ading or retrieval.<\/li>\n\n\n\n<li><strong>Poor train\u00ading data teach\u00ades unde\u00adsir\u00adable behav\u00adiours. <\/strong>The mod\u00adel faith\u00adful\u00adly copies mis\u00adtakes, bias\u00ades, and bad for\u00admat\u00adting in its exam\u00adples.<\/li>\n\n\n\n<li><strong>Over-spe\u00adcial\u00adi\u00adsa\u00adtion can reduce gen\u00ader\u00adal abil\u00adi\u00adty. <\/strong>Tun\u00ading too nar\u00adrow\u00adly can make a mod\u00adel worse at tasks out\u00adside the train\u00ading set.<\/li>\n\n\n\n<li><strong>More data is not auto\u00admat\u00adi\u00adcal\u00adly bet\u00adter. <\/strong>A small\u00ader, clean\u00ader, more diverse set often beats a large, noisy one.<\/li>\n\n\n\n<li><strong>Eval\u00adu\u00ada\u00adtion is required. <\/strong>With\u00adout a held-out test set you can\u00adnot con\u00adfirm that tun\u00ading actu\u00adal\u00adly helped.<\/li>\n\n\n\n<li><strong>It does not elim\u00adi\u00adnate hal\u00adlu\u00adci\u00adna\u00adtions. <\/strong>The mod\u00adel can still pro\u00adduce con\u00adfi\u00addent, incor\u00adrect answers.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Under\u00adstand\u00ading these lim\u00adits is part of answer\u00ading what is instruc\u00adtion tun\u00ading hon\u00adest\u00adly, and it sep\u00ada\u00adrates a real\u00adis\u00adtic plan from an over-opti\u00admistic one, which is why dis\u00adci\u00adplined eval\u00adu\u00ada\u00adtion is built into every seri\u00adous project.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Instruction tuning datasets<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Because the mod\u00adel imi\u00adtates its train\u00ading data, the dataset is the prod\u00aduct. Instruc\u00adtion tun\u00ading depends on sev\u00ader\u00adal kinds of data:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Instruc\u00adtion-response datasets: <\/strong>the core (instruc\u00adtion, input, out\u00adput) exam\u00adples used for super\u00advised fine-tun\u00ading.<\/li>\n\n\n\n<li><strong>Human-writ\u00adten demon\u00adstra\u00adtions: <\/strong>expert answers that show the mod\u00adel exact\u00adly how to respond.<\/li>\n\n\n\n<li><strong>Domain-spe\u00adcif\u00adic instruc\u00adtion data: <\/strong>exam\u00adples in your field (clin\u00adi\u00adcal, legal, code) for spe\u00adcialised assis\u00adtants.<\/li>\n\n\n\n<li><strong>RLHF pref\u00ader\u00adence datasets: <\/strong>rank\u00adings of com\u00adpet\u00ading respons\u00ades, used in the lat\u00ader align\u00adment stage.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Instruc\u00adtion tun\u00ading is not lim\u00adit\u00aded to text, either. Mul\u00adti\u00admodal mod\u00adels are instruc\u00adtion-tuned on image\u2011, audio\u2011, or video-and-response pairs, and first-per\u00adson sources such as <a href=\"https:\/\/www.graveiensai.com\/egocentric-video-data-collection\">ego\u00adcen\u00adtric video data<\/a> are increas\u00ading\u00adly used to teach mod\u00adels to fol\u00adlow instruc\u00adtions ground\u00aded in what a per\u00adson sees and does.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It also helps to think about the types of instruc\u00adtion data by what each one teach\u00ades the mod\u00adel:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th><strong>Data type<\/strong><\/th><th><strong>Pur\u00adpose<\/strong><\/th><\/tr><\/thead><tbody><tr><td><strong>Sin\u00adgle-turn instruc\u00adtions<\/strong><\/td><td>Basic instruc\u00adtion fol\u00adlow\u00ading<\/td><\/tr><tr><td><strong>Mul\u00adti-turn con\u00adver\u00adsa\u00adtions<\/strong><\/td><td>Con\u00adtext reten\u00adtion and dia\u00adlogue<\/td><\/tr><tr><td><strong>Domain-spe\u00adcif\u00adic instruc\u00adtions<\/strong><\/td><td>Spe\u00adcialised knowl\u00adedge and tasks<\/td><\/tr><tr><td><strong>Rea\u00adson\u00ading exam\u00adples<\/strong><\/td><td>Struc\u00adtured, step-by-step prob\u00adlem solv\u00ading<\/td><\/tr><tr><td><strong>Safe\u00adty exam\u00adples<\/strong><\/td><td>Appro\u00adpri\u00adate refusal and safe behav\u00adiour<\/td><\/tr><tr><td><strong>Pref\u00ader\u00adence pairs<\/strong><\/td><td>Lat\u00ader pref\u00ader\u00adence opti\u00admi\u00adsa\u00adtion (RLHF)<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">A strong instruc\u00adtion-tun\u00ading dataset shares a few prop\u00ader\u00adties, in this order of pri\u00ador\u00adi\u00adty:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>1. <\/strong><strong>Diver\u00adsi\u00adty: <\/strong>many task types, phras\u00adings, and domains, so the mod\u00adel gen\u00ader\u00adalis\u00ades.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>2. <\/strong><strong>Qual\u00adi\u00adty: <\/strong>accu\u00adrate, well-writ\u00adten tar\u00adget respons\u00ades, reviewed by qual\u00adi\u00adfied peo\u00adple.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>3. <\/strong><strong>Con\u00adsis\u00adten\u00adcy: <\/strong>a clear for\u00admat and labelling stan\u00addard applied across every exam\u00adple.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>4. <\/strong><strong>Cov\u00ader\u00adage: <\/strong>exam\u00adples for the edge cas\u00ades and behav\u00adiours you actu\u00adal\u00adly care about.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>5. <\/strong><strong>Safe\u00adty: <\/strong>exam\u00adples that teach the mod\u00adel to han\u00addle sen\u00adsi\u00adtive requests appro\u00adpri\u00adate\u00adly.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Assem\u00adbling this requires care\u00adful <a href=\"https:\/\/www.graveiensai.com\/data-collection\">data col\u00adlec\u00adtion<\/a>, skilled writ\u00aders and anno\u00adta\u00adtors, a clear guide\u00adline, and a mul\u00adti-stage qual\u00adi\u00adty-assur\u00adance and dataset-eval\u00adu\u00ada\u00adtion process. At <a href=\"https:\/\/www.graveiensai.com\/\">Graveiens AI<\/a>, we deliv\u00ader con\u00adsent-backed data, expert-writ\u00adten demon\u00adstra\u00adtions, RLHF pref\u00ader\u00adence data, and rig\u00ador\u00adous eval\u00adu\u00ada\u00adtion, through a four-stage work\u00adflow cer\u00adti\u00adfied to ISO 9001:2017. See <a href=\"https:\/\/www.graveiensai.com\/process\">how our process works<\/a> or read <a href=\"https:\/\/www.graveiensai.com\/why-choose-us\">why AI teams choose Graveiens AI<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Common challenges and best practices<\/strong><\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Over\u00adfit\u00adting to nar\u00adrow data: <\/strong>too few task types means the mod\u00adel fol\u00adlows instruc\u00adtions only in those pat\u00adterns. Fix it with delib\u00ader\u00adate diver\u00adsi\u00adty.<\/li>\n\n\n\n<li><strong>Incon\u00adsis\u00adtent labelling: <\/strong>mixed for\u00admats teach mixed sig\u00adnals. Fix it with a clear guide\u00adline and review, sup\u00adport\u00aded by strong <a href=\"https:\/\/www.graveiensai.com\/nlp\">NLP<\/a> exper\u00adtise.<\/li>\n\n\n\n<li><strong>Qual\u00adi\u00adty over quan\u00adti\u00adty: <\/strong>a small\u00ader set of excel\u00adlent, diverse exam\u00adples usu\u00adal\u00adly beats a huge, noisy one.<\/li>\n\n\n\n<li><strong>Weak eval\u00adu\u00ada\u00adtion: <\/strong>with\u00adout a held-out test set, you can\u00adnot tell whether tun\u00ading helped. Always mea\u00adsure against unseen instruc\u00adtions.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The through-line is sim\u00adple: the dis\u00adci\u00adpline of the dataset deter\u00admines the qual\u00adi\u00adty of the mod\u00adel. Teams that treat instruc\u00adtion tun\u00ading as a data prob\u00adlem, not just a train\u00ading run, get far bet\u00adter results.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Frequently asked questions<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Q. <\/strong><strong>What is instruc\u00adtion tun\u00ading in sim\u00adple terms?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A. <\/strong>A form of super\u00advised fine-tun\u00ading that trains a pre\u00adtrained lan\u00adguage mod\u00adel on exam\u00adples of instruc\u00adtions paired with cor\u00adrect respons\u00ades, so it learns to fol\u00adlow nat\u00adur\u00adal-lan\u00adguage com\u00admands. It turns a raw text pre\u00addic\u00adtor into a help\u00adful assis\u00adtant.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Q. <\/strong><strong>Is instruc\u00adtion tun\u00ading the same as super\u00advised fine-tun\u00ading?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A. <\/strong>They are close\u00adly relat\u00aded and often used inter\u00adchange\u00adably. Super\u00advised fine-tun\u00ading is the gen\u00ader\u00adal method of train\u00ading on labelled input-out\u00adput pairs; instruc\u00adtion tun\u00ading is SFT done with instruc\u00adtion-for\u00admat\u00adted data across many tasks.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Q. <\/strong><strong>What is the dif\u00adfer\u00adence between instruc\u00adtion tun\u00ading and RLHF?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A. <\/strong>Instruc\u00adtion tun\u00ading teach\u00ades a mod\u00adel to fol\u00adlow instruc\u00adtions using demon\u00adstra\u00adtions of cor\u00adrect respons\u00ades. RLHF and oth\u00ader pref\u00ader\u00adence-opti\u00admiza\u00adtion meth\u00adods then refine the mod\u00adel using human pref\u00ader\u00adence data. Instruc\u00adtion tun\u00ading usu\u00adal\u00adly comes first.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Q. <\/strong><strong>What is the dif\u00adfer\u00adence between instruc\u00adtion tun\u00ading and fine-tun\u00ading?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A. <\/strong>Fine-tun\u00ading is any fur\u00adther train\u00ading of a pre\u00adtrained mod\u00adel. Instruc\u00adtion tun\u00ading is a spe\u00adcif\u00adic kind of fine-tun\u00ading that uses instruc\u00adtion-for\u00admat\u00adted data across many tasks to teach gen\u00ader\u00adal instruc\u00adtion-fol\u00adlow\u00ading, rather than adapt\u00ading to one nar\u00adrow task.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Q. <\/strong><strong>What data do you need for instruc\u00adtion tun\u00ading?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A. <\/strong>A diverse, high-qual\u00adi\u00adty dataset of instruc\u00adtion-and-response pairs, writ\u00adten and reviewed by skilled peo\u00adple, with con\u00adsis\u00adtent for\u00admat\u00adting and cov\u00ader\u00adage of the behav\u00adiours and edge cas\u00ades you care about.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Q. <\/strong><strong>Can I instruc\u00adtion-tune an open-source mod\u00adel?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A. <\/strong>Yes. Open mod\u00adels such as Lla\u00adma can be instruc\u00adtion-tuned on your own dataset, often effi\u00adcient\u00adly with LoRA or oth\u00ader PEFT meth\u00adods. Qual\u00adi\u00adty depends almost entire\u00adly on the instruc\u00adtion data you pro\u00advide.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Q. <\/strong><strong>Does instruc\u00adtion tun\u00ading replace prompt engi\u00adneer\u00ading?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A. <\/strong>No. Instruc\u00adtion tun\u00ading changes the mod\u00adel at train\u00ading time; prompt engi\u00adneer\u00ading shapes behav\u00adiour at run time. They are com\u00adple\u00admen\u00adtary, and well-tuned mod\u00adels still ben\u00ade\u00adfit from good prompts.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Conclusion<\/strong><\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">So, what is instruc\u00adtion tun\u00ading? It is a form of super\u00advised fine-tun\u00ading that trains a pre\u00adtrained mod\u00adel on instruc\u00adtion-and-response pairs so it learns to fol\u00adlow com\u00admands. It sits between pre\u00adtrain\u00ading and align\u00adment meth\u00adods like RLHF, and it is the moment a raw lan\u00adguage mod\u00adel becomes a use\u00adful assis\u00adtant.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instruc\u00adtion tun\u00ading is one of the most impor\u00adtant post-train\u00ading tech\u00adniques for turn\u00ading a pre\u00adtrained lan\u00adguage mod\u00adel into an instruc\u00adtion-fol\u00adlow\u00ading sys\u00adtem. By train\u00ading on diverse, high-qual\u00adi\u00adty instruc\u00adtion-response data, teams can improve usabil\u00adi\u00adty, con\u00adsis\u00adten\u00adcy, and task per\u00adfor\u00admance with\u00adout train\u00ading a foun\u00adda\u00adtion mod\u00adel from scratch. The deep\u00ader les\u00adson is con\u00adsis\u00adtent across all of AI: the mod\u00adel learns from its data, which is why the col\u00adlec\u00adtion, writ\u00ading, and review of high-qual\u00adi\u00adty instruc\u00adtion data is the real deter\u00admi\u00adnant of suc\u00adcess.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Build\u00ading or fine-tun\u00ading an AI mod\u00adel?<\/strong>Talk to the Graveiens AI team about the instruc\u00adtion-tun\u00ading, super\u00advised fine-tun\u00ading, and RLHF datasets behind mod\u00adels that fol\u00adlow instruc\u00adtions reli\u00adably.&nbsp; <a href=\"https:\/\/www.graveiensai.com\/contact-us\"><strong>graveiensai.com\/contact-us<\/strong><\/a><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Sources: <\/em><a href=\"https:\/\/rlhfbook.com\/c\/04-instruction-tuning\" target=\"_blank\" rel=\"noopener\">Nathan Lam\u00adbert, RLHF Book (instruc\u00adtion fine-tun\u00ading)<\/a><em>; <\/em><a href=\"https:\/\/arxiv.org\/abs\/2109.01652\" target=\"_blank\" rel=\"noopener\">Wei et al., FLAN (arX\u00adiv)<\/a><em>; <\/em><a href=\"https:\/\/arxiv.org\/abs\/2203.02155\" target=\"_blank\" rel=\"noopener\">Ouyang et al., Instruct\u00adG\u00adPT (arX\u00adiv)<\/a><em>; <\/em><a href=\"https:\/\/www.ibm.com\/think\/topics\/instruction-tuning\" target=\"_blank\" rel=\"noopener\">IBM, instruc\u00adtion tun\u00ading explained<\/a><em>.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Instruc\u00adtion tun\u00ading is a form of super\u00advised fine-tun\u00ading that trains a pre\u00adtrained lan\u00adguage mod\u00adel on instruc\u00ad\u00adtion-and-response exam\u00adples so it learns to fol\u00adlow nat\u00adur\u00adal-lan\u00adguage com\u00admands. It teach\u00ades the mod\u00adel\u2026<\/p>\n","protected":false},"author":1,"featured_media":106,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"wp_typography_post_enhancements_disabled":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-105","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog"],"_links":{"self":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts\/105","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/comments?post=105"}],"version-history":[{"count":1,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts\/105\/revisions"}],"predecessor-version":[{"id":107,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/posts\/105\/revisions\/107"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/media\/106"}],"wp:attachment":[{"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/media?parent=105"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/categories?post=105"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.graveiensai.com\/blog\/wp-json\/wp\/v2\/tags?post=105"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}