文档教程【免费下载链接】learnxinyminutes-docsCode documentation written as code! How novel and totally my idea!项目地址https://gitcode.com/gh_mirrors/le/learnxinyminutes-docs点击查看免费下载本文基于 Learn X in Y minuteslearnxinyminutes-docs仓库中的 openmp.md以该文档的完整脉络为主线系统讲解 OpenMP 共享内存并行编程从 Master/Slave 线程模型与编译开关到omp.h运行时 API、私有/共享变量声明、七类同步指令再到循环并行化、加速比实测与分形计算实战。读完本文你可以独立完成 OpenMP 程序的编写、编译、线程数调优与并行加速验证。一、OpenMP 是什么OpenMP 是用于共享内存机器的并行编程库。它允许你使用简单的高层语法结构pragma 指令来表达并行性同时把底层线程管理的细节隐藏起来做到易用、编写快速。OpenMP 支持 C、C 和 Fortran 三种语言。在这个文档仓库的定位体系中OpenMP 并不是编程语言而是被归类为工具类教程。可以从 openmp.md 开头的 frontmatter 直接验证这一点--- category: tool name: OpenMP filename: learnopenMP.cpp contributors: - [Cillian Smith, https://github.com/smithc36-tcd] ---对照 CONTRIBUTING.md 的规范category字段可取language、tool或Algorithms Data Structures三类缺省为languagefilename字段指定了本篇教程代码的汇总文件名站点构建时会据此把正文中的代码合并成一个可下载的文件即learnopenMP.cpp。而 lint/frontmatter.py 中的allowed_keys校验列表name、category、filename、contributors、translators等保证了这类 frontmatter 在仓库层面通过格式检查。二、线程模型Master 与 Slave一个典型的 OpenMP 程序遵循如下结构引自 openmp.mdMaster主线程负责启动环境、初始化变量Slave从线程为被特殊指令directive标记的代码段而创建真正执行并行部分的线程。原文档用一张 ASCII 图刻画了主从结构__________ Slave /__________ Slave / Master ------------- Master \___________ Slave \__________ Slave每个线程都有唯一的 ID可通过omp_get_thread_num()获取下文会详细展开。三、第一个 OpenMP 程序编译与运行最简单的 hello world 程序可以通过#pragma omp parallel指令并行化#include stdio.h int main() { #pragma omp parallel { printf(Hello, World!\n); } return 0; }关键点在于编译时必须打开 OpenMP 开关而开关取决于具体编译器这是原文档明确提示的编译器编译参数Intel-openmpGCC-fopenmppgcc-mp以 GCC 为例gcc -fopenmp hello.c -o Hello运行后输出形如Hello, World! ... Hello, World!Hello, World! 的打印次数取决于机器的核心数原文档作者在自己的笔记本上得到了 12 次。换句话说不加任何控制时omp parallel区域默认会派生出与处理器核数相匹配的一组线程每个线程各打印一次。四、线程控制与omp.h运行时 API4.1 用环境变量控制线程数最直接的方式是设置环境变量export OMP_NUM_THREADS8这样即可把默认线程数改为 8无需修改任何代码。4.2omp.h中的常用库函数omp.h头文件提供了一组运行时函数用于查询和调整线程状态完整清单来自 openmp.md// Check the number of threads printf(Max Threads: %d\n, omp_get_max_threads()); printf(Current number of threads: %d\n, omp_get_num_threads()); printf(Current Thread ID: %d\n, omp_get_thread_num()); // Modify the number of threads omp_set_num_threads(int); // Check if we are in a parallel region omp_in_parallel(); // Dynamically vary the number of threads omp_set_dynamic(int); omp_get_dynamic(); // Check the number of processors printf(Number of processors: %d\n, omp_num_procs());这些函数按用途可分为三类整理如下函数作用类别omp_get_max_threads()查询当前最大线程数查询omp_get_num_threads()查询当前所在并行区域内线程数查询omp_get_thread_num()查询当前线程 ID查询omp_num_procs()查询处理器核数量查询omp_in_parallel()判断当前是否处于并行区域内查询omp_set_num_threads(int)设置线程数控制omp_set_dynamic(int)/omp_get_dynamic()开关/查询动态调整线程数行为控制其中omp_in_parallel()常用于区分主线程与从线程的执行路径omp_set_dynamic与omp_get_dynamic构成一对开关接口允许运行时动态伸缩线程数量原文档将其归纳为 Dynamically vary the number of threads。五、私有变量与共享变量并行区域内变量可以被声明为私有private或共享shared原文示例见 openmp.md// Variables in parallel sections can be either private or shared. /* Private variables are private to each thread, as each thread has its own * private copy. These variables are not initialized or maintained outside * the thread. */ #pragma omp parallel private(x, y) /* Shared variables are visible and accessible by all threads. By default, * all variables in the work sharing region are shared except the loop * iteration counter. * * Shared variables should be used with care as they can cause race conditions. */ #pragma omp parallel shared(a, b, c) // They can be declared together as follows #pragma omp parallel private(x, y) shared(a,b,c)要点归纳私有变量每个线程持有独立副本互不干扰线程之外不初始化、不维护生命周期只存在于该线程内。私有变量天然不存在竞争。共享变量所有线程可见、可访问。默认规则是——除循环迭代计数器外工作共享区域中的变量默认都是共享的因此循环索引不会与外层变量冲突但其他全局/外部变量都会。共享变量必须小心使用因为它们可能引发竞态条件race condition需要配合第六节的同步指令来保护。private与shared子句可以在同一条指令中并列书写。六、同步指令critical、single、atomic、ordered、barrier、nowait、reductionOpenMP 提供了一组指令来控制线程同步。原文档在同一个并行区域内集中演示了全部七种写法完整代码见 openmp.md#pragma omp parallel { /* critical: the enclosed code block will be executed by only one thread * at a time, and not simultaneously executed by multiple threads. It is * often used to protect shared data from race conditions. */ #pragma omp critical data data computed; /* single: used when a block of code needs to be run by only a single * thread in a parallel section. Good for managing control variables. */ #pragma omp single printf(Current number of threads: %d\n, omp_get_num_threads()); /* atomic: Ensures that a specific memory location is updated atomically * to avoid race conditions. */ #pragma omp atomic counter 1; /* ordered: the structured block is executed in the order in which * iterations would be executed in a sequential loop */ #pragma omp for ordered for (int i 0; i N; i) { #pragma omp ordered process(data[i]); } /* barrier: Forces all threads to wait until all threads reach this point * before proceeding. */ #pragma omp barrier /* nowait: Allows threads to proceed with their next task without waiting * for other threads to complete the current one. */ #pragma omp for nowait for (int i 0; i N; i) { process(data[i]); } /* reduction : Combines the results of each threads computation into a * single result. */ #pragma omp parallel for reduction(:sum) for (int i 0; i N; i) { sum a[i] * b[i]; } }按语义逐条拆解指令作用典型用途critical被包围的代码块同一时刻只允许一个线程执行保护共享数据免遭竞态是通用的互斥手段single并行区域内仅由一个线程执行该代码块管理控制变量、避免重复输出如只打印一次状态atomic保证对某一内存位置的更新是原子的简单计数器等轻量共享更新比critical开销更小ordered结构化块按顺序循环的执行次序依次执行需与for联用块内以#pragma omp ordered标记结果必须按迭代顺序产生/输出的场景barrier强制所有线程等待直到全部线程到达该点后才继续阶段间同步nowait线程无需等待其他线程完成当前工作即可进入下一任务消除工作共享构造末尾的隐式屏障降低同步开销reduction把每个线程的局部计算结果归约为单一结果如reduction(:sum)求和、求最值等并行归约barrier 完整示例原文档给出了一个最小可运行的barrier示例openmp.md它同时演示了omp_get_num_threads()、omp_get_thread_num()两个 API 的用法#include omp.h #include stdio.h int main() { // Current number of active threads printf(Num of threads is %d\n, omp_get_num_threads()); #pragma omp parallel { // Current thread ID printf(Thread ID: %d\n, omp_get_thread_num()); #pragma omp barrier --- Wait here until other threads have returned if(omp_get_thread_num() 0) { printf(\nNumber of active threads: %d\n, omp_get_num_threads()); } } return 0; }注意#pragma omp barrier后面那句Wait here until other threads have returned是原文档对指令位置的标注不是代码的一部分实际编译时应去掉。示例逻辑是主线程先打印当前线程数进入并行区域后每个线程打印自己的 IDbarrier保证所有线程都打印完 ID 之后才由 0 号线程打印活动线程总数。七、循环并行化与数据依赖约束用工作共享work-sharing指令并行化循环非常直接#pragma omp parallel { #pragma omp for // for loop to be parallelized for() ... }必须强调原文档给出的前提条件openmp.md循环必须易于并行化OpenMP 才能将其展开并在各线程间分配迭代如果相邻迭代之间存在数据依赖data dependenciesOpenMP 无法并行化该循环——因为某次迭代的结果会依赖前一次迭代尚未完成的结果强行拆分会破坏正确性。八、串行 vs 并行加速比实测原文档提供了一个 C 程序在同一台机器上对比串行循环与#pragma omp parallel for的执行时间完整代码见 openmp.md#include iostream #include vector #include ctime #include chrono #include omp.h int main() { const int num_elements 1e8; std::vectordouble a(num_elements, 1.0); std::vectordouble b(num_elements, 2.0); std::vectordouble c(num_elements, 0.0); // Serial version auto start_time std::chrono::high_resolution_clock::now(); for (int i 0; i num_elements; i) { c[i] a[i] * b[i]; } auto end_time std::chrono::high_resolution_clock::now(); auto duration_serial std::chrono::duration_caststd::chrono::milliseconds(end_time - start_time).count(); // Parallel version with OpenMP start_time std::chrono::high_resolution_clock::now(); #pragma omp parallel for for (int i 0; i num_elements; i) { c[i] a[i] * b[i]; } end_time std::chrono::high_resolution_clock::now(); auto duration_parallel std::chrono::duration_caststd::chrono::milliseconds(end_time - start_time).count(); std::cout Serial execution time: duration_serial ms std::endl; std::cout Parallel execution time: duration_parallel ms std::endl; std::cout Speedup: static_castdouble(duration_serial) / duration_parallel std::endl; return 0; }注原文代码写作100000000即一亿个元素上面保持数值等价三组std::vectordouble分别初始化为 1.0 / 2.0 / 0.0。原文档记录的运行结果为Serial execution time: 488 ms Parallel execution time: 148 ms Speedup: 3.2973原文档同时给出了两点诚实的保留意见做性能评估时同样应当记住该例相当构造化contrived真实加速比取决于具体实现由于缓存性能cache performance等原因串行代码有时反而可能比并行代码更快——并行化并非无条件加速omp parallel for的调度、线程间通信和缓存行为都会影响最终表现。九、完整实战OpenMP 计算 Mandelbrot 集合原文档的收官示例用 OpenMP 计算 2000×2000 的 Mandelbrot 集合并输出 PPM 图像完整代码见 openmp.md#include iostream #include fstream #include complex #include vector #include omp.h const int width 2000; const int height 2000; const int max_iterations 1000; int mandelbrot(const std::complexdouble c) { std::complexdouble z c; int n 0; while (abs(z) 2 n max_iterations) { z z * z c; n; } return n; } int main() { std::vectorstd::vectorint values(height, std::vectorint(width)); // Calculate the Mandelbrot set using OpenMP #pragma omp parallel for schedule(dynamic) for (int y 0; y height; y) { for (int x 0; x width; x) { double real (x - width / 2.0) * 4.0 / width; double imag (y - height / 2.0) * 4.0 / height; std::complexdouble c(real, imag); values[y][x] mandelbrot(c); } } // Prepare the output image std::ofstream image(mandelbrot_set.ppm); image P3\n width height 255\n; // Write the output image in serial for (int y 0; y height; y) { for (int x 0; x width; x) { int value values[y][x]; int r (value % 8) * 32; int g (value % 16) * 16; int b (value % 32) * 8; image r g b ; } image \n; } image.close(); std::cout Mandelbrot set image generated as mandelbrot_set.ppm. std::endl; return 0; }该示例集中体现了前文多个知识点值得逐点对照并行化对象的选择只有最外层按行y的循环被#pragma omp parallel for并行化。各行之间没有数据依赖每行的values[y][x]只写自己那一行满足第七节的可并行化前提行内的x循环则保持串行。schedule(dynamic)子句Mandelbrot 各行迭代代价并不均匀——靠近分形边界的像素迭代次数接近max_iterations 1000外部点则很快逃逸。动态调度让先完成空闲的线程从剩余行中动态领取工作从源码结构看这是针对各行计算量不均的合理选择可以推断静态均分会造成线程负载不均衡。结果写入保持串行输出阶段用单个std::ofstream顺序写 PPM 文件避免了多线程同时写同一个输出流的同步问题——这也呼应了共享资源要保护/串行化的原则。图像格式文件头P3\n2000 2000 255\n是 PPM 文本格式的头部宽 高 最大值随后逐像素写 R、G、B配色由迭代次数value取模生成value越大的区域颜色变化越丰富。产物运行后生成mandelbrot_set.ppm并在控制台提示生成完成。十、回到仓库这篇文档在 learnxinyminutes-docs 中的组织方式把视角拉回文档仓库本身openmp.md 的组织方式体现了该项目的两条核心规范见 CONTRIBUTING.mdFrontmatter 驱动站点生成category: tool决定其归类filename: learnopenMP.cpp决定正文代码如何被合并导出为可下载的单一文件contributors用于署名。这些字段会被 lint/frontmatter.py 的process_files()逐一校验YAML 格式 白名单键名 值类型。风格约定行宽尽量控制在 80 字符内、示例优先于论述Prefer example to exposition、面向有经验的程序员保持简洁——这正是本文各章节大量采用代码块 简短批注形式的原因。站点构建项目文档在 README.md 与 CONTRIBUTING.md 中说明了本地构建方式——克隆站点仓库与本仓库、pip install -r requirements.txt、python build.py再用python -m http.server预览本仓库只保存 Markdown 源文档与 lint 工具见 lint/ 目录含 lint/encoding.sh、lint/requirements.txt。原文档末尾还列出了若干外部参考资源并行编程入门讲义、OpenMP 官方教程合集、OpenMP 官方速查卡。受限于本文不外链的约束这里不再列出感兴趣的读者可对照 openmp.md 末尾的 Resources 一节自行查阅。总结沿着 openmp.md 的脉络本文完成了从入门到实战的完整闭环模型Master 负责初始化Slave 执行并行区域线程 ID 由omp_get_thread_num()获取工具链GCC 用-fopenmp、Intel 用-openmp、pgcc 用-mp开启支持export OMP_NUM_THREADSN控制线程数APIomp_get_max_threads/omp_get_num_threads/omp_num_procs/omp_in_parallel/omp_set_num_threads/omp_set_dynamic覆盖查询与控制变量语义默认共享循环迭代计数器除外显式private/shared子句消除歧义共享变量需防竞态同步critical互斥、single单线程执行、atomic原子更新、ordered有序执行、barrier屏障、nowait免等待、reduction归约各司其职实战可并行化循环的判定标准无跨迭代数据依赖、加速比实测的方法与缓存性能带来的反直觉结论、Mandelbrot 示例中schedule(dynamic)对负载不均的应对。这些内容既可以直接作为日常开发的 OpenMP 速查手册也为深入理解共享内存并行的同步与调度机制提供了起点。赞分享文档教程【免费下载链接】learnxinyminutes-docsCode documentation written as code! How novel and totally my idea!项目地址https://gitcode.com/gh_mirrors/le/learnxinyminutes-docs点击查看免费下载相关推荐从Python到RustLearn X in Y Minutes教程风格解析从Python到RustLearn X in Y Minutes教程风格解析 本文深入分析了Learn X in Y Minutes项目中不同编程语言教程的教文档教程几分钟看懂 jQuery选择器、事件、效果与 DOM 操作Learn X in Y Minutes 风格解析几分钟看懂 jQuery选择器、事件、效果与 DOM 操作Learn X in Y Minutes 风格解析 本文基于 Learn X in Y Minu文档教程Learn X in Y minutes意大利语版 C 入门教程全解析Learn X in Y minutes意大利语版 C 入门教程全解析 本文基于本仓库 it/c.md https://link.gitcode.co文档教程上一篇Vrite看板管理教程从内容创建到发布的完整流程下一篇5分钟掌握Trippy网络诊断新利器完全指南创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
阅读完成 · 觉得有帮助?