文档教程提示工程大模型人工智能RAGAI Agent【免费下载链接】Prompt-Engineering-Guide Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.项目地址https://gitcode.com/GitHub_Trending/pr/Prompt-Engineering-Guide点击查看免费下载本文是 Prompt-Engineering-Guide 仓库中「Advanced Prompting」系列指南guides/prompts-advanced-usage.md的技术解读与实践手册系统梳理了零样本提示、少样本提示、思维链CoT、零样本 CoT、自洽性Self-Consistency、生成知识提示与自动提示工程师APE七大进阶技术。读完本文你将理解每种技术的适用场景、完整提示模板与输出效果并能结合仓库中的双语 Web 版本pages/techniques 目录下各语言*.en.mdx文档与 Notebook 动手验证构建一套可复制的进阶提示策略。为什么需要更高级的提示技术在基础指南guides/prompts-basic-usage.md中我们展示了文本摘要、信息抽取、问答、分类、对话、代码生成与基础推理等任务。随着任务复杂度的上升仅仅给出指令往往不够——模型在需要多步推理的任务上会频繁出错。本文要解决的核心问题就是当零样本与少样本提示失效时如何用更系统的提示策略让 LLM 完成算术、常识与符号推理等复杂任务。下文将按从简单到高级的顺序逐一展开这些技术。零样本提示Zero-Shot Prompting现代 LLM如 GPT-3.5 Turbo、GPT-4、Claude 3经过大规模训练并针对指令进行了调优天然具备**零样本zero-shot**执行任务的能力——即提示中不包含任何示例或演示demonstration模型直接根据指令完成任务。最典型的零样本示例是情感分类Classify the text into neutral, negative, or positive. Text: I think the vacation is okay. Sentiment:模型输出Neutral注意上面的提示没有提供任何文本—标签配对样例模型却能输出正确情感这正是零样本能力的体现。从仓库的 pages/techniques/zeroshot.en.mdx 可知这一能力背后依赖两类关键技术指令微调instruction tuning在由指令描述的数据集上对模型进行微调显著提升零样本学习表现源自 Wei et al. 2022 所引的指令微调工作基于人类反馈的强化学习RLHF将指令微调进一步对齐到人类偏好这正是 ChatGPT 类模型的基石。适用判断当零样本提示失效模型给出错误或不稳定结果时应优先尝试在提示中补充演示或示例即进入少样本提示阶段。少样本提示Few-Shot Prompting少样本提示通过**上下文学习in-context learning**在提示中提供若干演示让演示作为后续输入的条件conditioning引导模型生成更符合预期的响应。这是应对复杂任务最直接的增强手段。1-shot 演示让模型学会造词造句以下示例来自 Brown et al. 2020GPT-3 论文任务是根据定义正确使用一个新造的词A whatpu is a small, furry animal native to Tanzania. An example of a sentence that uses the word whatpu is: We were traveling in Africa and we saw these very cute whatpus. To do a farduddle means to jump up and down really fast. An example of a sentence that uses the word farduddle is:模型输出When we won the game, we all started to farduddle in celebration.只给 1 个示例1-shot模型就掌握了规律。对更困难的任务可以实验性地增加演示数量3-shot、5-shot、10-shot 等观察性能变化。设计演示的关键要点Min et al. 2022Min et al. (2022) 的研究给出三条关于演示设计的结论值得在构建少样本提示时反复对照标签空间与输入文本分布都很重要——无论单个输入的标签是否正确演示所规定的标签空间与输入文本分布都会显著影响性能格式本身影响性能——即使使用随机标签也比完全没有标签好得多标签采样方式有讲究——从标签的真实分布中采样随机标签而非均匀分布会带来额外收益。实验随机标签仍然有效下面的提示故意把 Negative/Positive 标签随机分配给输入This is awesome! // Negative This is bad! // Positive Wow that movie was rad! // Positive What a horrible show! //输出Negative尽管标签被随机化模型依然得到了正确答案——格式被保留是关键因素之一。更进一步较新的 GPT 系列模型对随机格式也表现出更强的鲁棒性例如Positive This is awesome! This is bad! Negative Wow that movie was rad! Positive What a horrible show! --输出依然是Negative但需要注意该结论尚未在更复杂、更多样的任务上得到系统验证实际应用时应自行测试。少样本提示的局限少样本提示在多步推理任务上并不稳定。以奇数之和是否为偶数的推理题为例直接零样本询问时模型会给出错误答案The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型错误输出Yes, the odd numbers in this group add up to 107, which is an even number.补充 4 个演示后只给出 True/False 结论不展示推理过程The odd numbers in this group add up to an even number: 4, 8, 9, 15, 12, 2, 1. A: The answer is False. The odd numbers in this group add up to an even number: 17, 10, 19, 4, 8, 12, 24. A: The answer is True. The odd numbers in this group add up to an even number: 16, 11, 14, 4, 8, 13, 24. A: The answer is True. The odd numbers in this group add up to an even number: 17, 9, 10, 12, 13, 4, 2. A: The answer is False. The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型输出The answer is True.答案依然错误正确应为 False。这说明该任务需要多个推理步骤而仅仅提供结论型演示不足以让模型学会推理路径。由此引出的结论是当零样本与少样本都不够用时需要把问题拆解成步骤并展示给模型——这正是思维链提示的出发点。思维链提示Chain-of-Thought Prompting思维链CoT由 Wei et al. (2022) 提出核心思想是通过中间推理步骤激发模型的复杂推理能力并可将其与少样本提示结合使用在需要先推理再回答的任务上取得更好效果。继续使用上面的推理题但这次在每个演示中展示完整的推理步骤The odd numbers in this group add up to an even number: 4, 8, 9, 15, 12, 2, 1. A: Adding all the odd numbers (9, 15, 1) gives 25. The answer is False. The odd numbers in this group add up to an even number: 17, 10, 19, 4, 8, 12, 24. A: Adding all the odd numbers (17, 19) gives 36. The answer is True. The odd numbers in this group add up to an even number: 16, 11, 14, 4, 8, 13, 24. A: Adding all the odd numbers (11, 13) gives 24. The answer is True. The odd numbers in this group add up to an even number: 17, 9, 10, 12, 13, 4, 2. A: Adding all the odd numbers (17, 9, 13) gives 39. The answer is False. The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型输出Adding all the odd numbers (15, 5, 13, 7, 1) gives 41. The answer is False.结果完全正确。更有趣的是即使只提供 1 个带推理步骤的演示也足够The odd numbers in this group add up to an even number: 4, 8, 9, 15, 12, 2, 1. A: Adding all the odd numbers (9, 15, 1) gives 25. The answer is False. The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1. A:模型输出Adding all the odd numbers (15, 5, 13, 7, 1) gives 41. The answer is False.需要说明的是论文作者指出这种能力属于涌现能力emergent ability只有在模型规模足够大时才会出现。仓库中完整的 CoT 讲解与演示图见 pages/techniques/cot.en.mdx。零样本思维链Zero-Shot CoTKojima et al. (2022) 提出了一种极简的零样本思维链方案在原提示末尾追加一句 Lets think step by step即可触发模型分步推理。这在没有足够示例可用的场景下尤其有价值。先看不加特殊提示的普通提问I went to the market and bought 10 apples. I gave 2 apples to the neighbor and 2 to the repairman. I then went and bought 5 more apples and ate 1. How many apples did I remain with?模型错误输出11 apples加上特殊提示后再试I went to the market and bought 10 apples. I gave 2 apples to the neighbor and 2 to the repairman. I then went and bought 5 more apples and ate 1. How many apples did I remain with? Lets think step by step.模型输出First, you started with 10 apples. You gave away 2 apples to the neighbor and 2 to the repairman, so you had 6 apples left. Then you bought 5 more apples, so now you had 11 apples. Finally, you ate 1 apple, so you would remain with 10 apples.仅仅一行提示就能把错误答案11纠正为正确答案10且模型完整展示了推理链条。这正是零样本 CoT 的魅力用最小的提示成本换取可解释的推理过程。自洽性Self-Consistency自洽性由 Wang et al. (2022) 提出目标是取代思维链提示中朴素的贪心解码通过少样本 CoT 采样出多条多样化的推理路径再基于这些生成结果选出最一致的答案从而在算术与常识推理任务上进一步提升 CoT 的性能。先看一个单次解码会出错的例子When I was 6 my sister was half my age. Now I’m 70 how old is my sister?模型错误输出35使用 Wang et al. 2022 论文Table 17中的少样本示例配合多次采样Q: There are 15 trees in the grove. Grove workers will plant trees in the grove today. After they are done, there will be 21 trees. How many trees did the grove workers plant today? A: We start with 15 trees. Later we have 21 trees. The difference must be the number of trees they planted. So, they must have planted 21 - 15 6 trees. The answer is 6. Q: If there are 3 cars in the parking lot and 2 more cars arrive, how many cars are in the parking lot? A: There are 3 cars in the parking lot already. 2 more arrive. Now there are 3 2 5 cars. The answer is 5. Q: Leah had 32 chocolates and her sister had 42. If they ate 35, how many pieces do they have left in total? A: Leah had 32 chocolates and Leah’s sister had 42. That means there were originally 32 42 74 chocolates. 35 have been eaten. So in total they still have 74 - 35 39 chocolates. The answer is 39. Q: Jason had 20 lollipops. He gave Denny some lollipops. Now Jason has 12 lollipops. How many lollipops did Jason give to Denny? A: Jason had 20 lollipops. Since he only has 12 now, he must have given the rest to Denny. The number of lollipops he has given to Denny must have been 20 - 12 8 lollipops. The answer is 8. Q: Shawn has five toys. For Christmas, he got two toys each from his mom and dad. How many toys does he have now? A: He has 5 toys. He got 2 from mom, so after that he has 5 2 7 toys. Then he got 2 more from dad, so in total he has 7 2 9 toys. The answer is 9. Q: There were nine computers in the server room. Five more computers were installed each day, from monday to thursday. How many computers are now in the server room? A: There are 4 days from monday to thursday. 5 computers were added each day. That means in total 4 * 5 20 computers were added. There were 9 computers in the beginning, so now there are 9 20 29 computers. The answer is 29. Q: Michael had 58 golf balls. On tuesday, he lost 23 golf balls. On wednesday, he lost 2 more. How many golf balls did he have at the end of wednesday? A: Michael initially had 58 balls. He lost 23 on Tuesday, so after that he has 58 - 23 35 balls. On Wednesday he lost 2 more so now he has 35 - 2 33 balls. The answer is 33. Q: Olivia has $23. She bought five bagels for $3 each. How much money does she have left? A: She bought 5 bagels for $3 each. This means she spent 5 Q: When I was 6 my sister was half my age. Now I’m 70 how old is my sister? A:多次采样得到的输出输出 1When I was 6 my sister was half my age, so she was 3. Now I am 70, so she is 70 - 3 67. The answer is 67.输出 2When the narrator was 6, his sister was half his age, which is 3. Now that the narrator is 70, his sister would be 70 - 3 67 years old. The answer is 67.输出 3When I was 6 my sister was half my age, so she was 3. Now I am 70, so she is 70/2 35. The answer is 35.可以看到前两条路径都得到67出现多数一致答案最终即可确定为 67。真实应用中求最终答案还需结合投票等步骤详见原论文但从工程角度理解其核心用多次采样 一致性投票对冲单次贪心解码的偶然错误。仓库中的完整演示见 pages/techniques/consistency.en.mdx。生成知识提示Generated Knowledge PromptingLLM 的一个常见短板是缺乏世界知识。生成知识提示的思路来自 Liu et al. 2022在模型做预测之前先让模型自己生成相关知识并把知识并入提示中再基于知识作答。该方法对常识推理类任务尤其有帮助。第一步看一个常识错误Part of golf is trying to get a higher point total than others. Yes or No?模型错误输出Yes.这个错误暴露出 LLM 对世界知识的理解局限——高尔夫恰恰是杆数越低越好。第二步先生成若干条知识提示模型为一系列输入生成知识Input: Greece is larger than mexico. Knowledge: Greece is approximately 131,957 sq km, while Mexico is approximately 1,964,375 sq km, making Mexico 1,389% larger than Greece. Input: Glasses always fog up. Knowledge: Condensation occurs on eyeglass lenses when water vapor from your sweat, breath, and ambient humidity lands on a cold surface, cools, and then changes into tiny drops of liquid, forming a film that you see as fog. Your lenses will be relatively cool compared to your breath, especially when the outside air is cold. Input: A fish is capable of thinking. Knowledge: Fish are more intelligent than they appear. In many areas, such as memory, their cognitive powers match or exceed those of ’higher’ vertebrates including non-human primates. Fish’s long-term memories help them keep track of complex social relationships. Input: A common effect of smoking lots of cigarettes in one’s lifetime is a higher than normal chance of getting lung cancer. Knowledge: Those who consistently averaged less than one cigarette per day over their lifetime had nine times the risk of dying from lung cancer than never smokers. Among people who smoked between one and 10 cigarettes per day, the risk of dying from lung cancer was nearly 12 times higher than that of never smokers. Input: A rock is the same size as a pebble. Knowledge: A pebble is a clast of rock with a particle size of 4 to 64 millimetres based on the Udden-Wentworth scale of sedimentology. Pebbles are generally considered larger than granules (2 to 4 millimetres diameter) and smaller than cobbles (64 to 256 millimetres diameter). Input: Part of golf is trying to get a higher point total than others. Knowledge:模型为最后一个输入生成两条候选知识知识 1The objective of golf is to play a set of holes in the least number of strokes. A round of golf typically consists of 18 holes. Each hole is played once in the round on a standard golf course. Each stroke is counted as one point, and the total number of strokes is used to determine the winner of the game.知识 2Golf is a precision club-and-ball sport in which competing players (or golfers) use many types of clubs to hit balls into a series of holes on a course using the fewest number of strokes. The goal is to complete the course with the lowest score, which is calculated by adding up the total number of strokes taken on each hole. The player with the lowest score wins the game.第三步整合知识做出预测将问题改写为 QA 格式并拼接知识引导模型给出带解释的答案Question: Part of golf is trying to get a higher point total than others. Yes or No? Knowledge: The objective of golf is to play a set of holes in the least number of strokes. A round of golf typically consists of 18 holes. Each hole is played once in the round on a standard golf course. Each stroke is counted as one point, and the total number of strokes is used to determine the winner of the game. Explain and Answer:答案 1置信度非常高No, the objective of golf is not to get a higher point total than others. Rather, the objective is to play a set of holes in the least number of strokes. The total number of strokes is used to determine the winner of the game, not the total number of points.换用知识 2 再试Question: Part of golf is trying to get a higher point total than others. Yes or No? Knowledge: Golf is a precision club-and-ball sport in which competing players (or golfers) use many types of clubs to hit balls into a series of holes on a course using the fewest number of strokes. The goal is to complete the course with the lowest score, which is calculated by adding up the total number of strokes taken on each hole. The player with the lowest score wins the game. Explain and Answer:答案 2置信度明显更低Yes, part of golf is trying to get a higher point total than others. Each player tries to complete the course with the lowest score, which is calculated by adding up the total number of strokes taken on each hole. The player with the lowest score wins the game.这个对比非常有意思不同知识会带来不同结论与置信度说明知识的质量直接决定预测质量。工程化落地时通常需要结合多条候选知识的答案与置信度做汇总详见原论文。本节的完整示例见仓库 pages/techniques/knowledge.en.mdx。自动提示工程师Automatic Prompt Engineer, APEZhou et al. (2022) 提出了 APE 框架用于自动化地生成与筛选指令把指令生成问题建模为自然语言综合任务并作为一个黑盒优化问题由 LLM 自身生成候选方案并搜索最优指令。APE 的两阶段流程候选指令生成由一个 LLM作为推理模型接收任务的输出演示output demonstrations生成一批指令候选候选指令评估与选择用目标模型执行这些指令基于计算出的评估分数evaluation scores选出最合适的指令。这一生成—执行—打分—选择的闭环把提示工程从纯人工试错提升为可自动化的搜索过程。APE 的经典成果发现优于人工的零样本 CoT 提示APE 发现了一个比人工设计的 Lets think step by stepKojima et al. 2022更好的零样本 CoT 提示Lets work this out in a step by step way to be sure we have the right answer.该提示能够触发思维链推理并在 MultiArith 与 GSM8K 基准上提升性能这直接证明了自动提示优化在真实基准上的价值——机器发现的提示可以超越人类专家手工编写的提示。相关自动优化方向若对提示自动优化感兴趣原文档与 pages/techniques/ape.en.mdx 还整理了以下代表性工作AutoPrompt基于梯度引导搜索为多样化任务自动创建提示Prefix Tuning微调的轻量替代方案为 NLG 任务前置一段可训练连续前缀Prompt Tuning通过反向传播学习软提示soft prompts。这些方法共同构成了提示优化这一重要研究分支是 APE 之外值得深入的方向。进阶路线与仓库配套资源掌握本文七项技术后可以按以下路径在仓库中继续深化对照 Web 版本精读本文内容在仓库中以多语言 MDX 页面组织英文版依次位于 pages/techniques/zeroshot.en.mdx、pages/techniques/fewshot.en.mdx、pages/techniques/cot.en.mdx、pages/techniques/consistency.en.mdx、pages/techniques/knowledge.en.mdx、pages/techniques/ape.en.mdx可对照中英文版本加深理解动手实践仓库提供了配套 Notebook notebooks/pe-lecture.ipynb可用于在真实 API 上复现本文各技术的提示模板与输出效果衔接前后章节本文是进阶部分向前承接 guides/prompts-basic-usage.md基础提示向后衔接 guides/prompts-applications.md提示工程应用本系列全部指南的索引见 guides/README.md注意适用范围本文各示例中的模型是否输出正确结果会随模型版本与参数设置如温度、采样次数变化尤其是自洽性与生成知识技术对采样数量、评估方式敏感落地前应基于目标模型自行验证。小结如何选择进阶提示策略根据任务类型可以将本文技术整理为如下决策顺序场景推荐技术关键操作模型能直接完成零样本提示仅提供清晰指令零样本失效少样本提示在提示中提供 25 条演示需要多步推理思维链CoT演示中展示推理步骤缺少示例零样本 CoT追加 Lets think step by step推理结果不稳定自洽性多次采样 一致性投票依赖世界知识生成知识提示先生成知识再基于知识作答提示优化成本高自动提示工程师APE让 LLM 自动生成并筛选指令从零样本到自动提示工程师这条进阶路径的底层逻辑始终一致不是模型本身变强了而是我们学会用更结构化的方式把推理路径、知识与评价标准暴露给模型。把这七项技术配合 notebooks/pe-lecture.ipynb 反复演练即可在真实业务中构建出稳定、可解释、可复用的高级提示方案。赞分享文档教程提示工程大模型人工智能RAGAI Agent【免费下载链接】Prompt-Engineering-Guide Guides, papers, lessons, notebooks and resources for prompt engineering, context engineering, RAG, and AI Agents.项目地址https://gitcode.com/GitHub_Trending/pr/Prompt-Engineering-Guide点击查看免费下载相关推荐ARIS 接入 OpenRouter用免费或按量模型搭建跨模型审稿后端的完整指南ARIS 接入 OpenRouter用免费或按量模型搭建跨模型审稿后端的完整指南 ARISAuto Research In Sleep以 MarkdownAI 技能/插件AI 评测科研人工智能MCP 服务dsh-plugin提示工程进阶掌握思维链技术的完整指南提示工程进阶掌握思维链技术的完整指南 欢迎来到Awesome Prompt Engineering项目中的思维链技术深度解析思维链Chain of Tho提示工程教程UI-TARS桌面版终极指南如何用AI自然语言控制你的电脑和浏览器UI TARS桌面版终极指南如何用AI自然语言控制你的电脑和浏览器 UI TARS desktop是一个基于视觉语言模型的开源多模态AI助手让你通过自然语人工智能大模型AI Agent桌面应用GUI 自动化浏览器控制MCP 服务MCP Clients上一篇告别手忙脚乱渔人的直感钓鱼计时器让FF14钓鱼从盯梢苦差变成指尖享受下一篇免费开源用OpenBoardView轻松查看.brd电路板文件创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
阅读完成 · 觉得有帮助?