NLP深度学习【免费下载链接】pytextA natural language modeling framework based on PyTorch项目地址https://gitcode.com/gh_mirrors/py/pytext点击查看免费下载导读本文基于 PyText 官方教程完整讲解如何在ATISAirline Travel Information System航班查询数据集上用 PyText 训练一个联合的意图检测Intent Detection与槽位填充Slot Filling模型从原始数据预处理、GloVe 预训练词向量加载到模型配置、训练调优再到用命令行对真实话语utterance生成预测。读完本文你将掌握 PyText 中IntentSlotTask/IntentSlotModel的完整使用链路并能对照仓库源码理解 BiLSTM Attention CRF 联合建模的内部原理。提示原教程标注为OBSOLETE声明其使用的是旧版 API。本文以当前仓库实际代码为准进行核对——IntentSlotTask仍注册在 pytext/builtin_task.py 中教程中的 JSON 配置结构与新版IntentSlotModelpytext/models/joint_model.py完全对应命令仍可执行但具体行为请以仓库当前版本为准。1. 任务背景为什么把意图检测与槽位填充做成联合模型在个人助理personal assistant的自然语言理解NLU场景中有两个最基础的任务意图检测Intent Detection判断用户这句话想干什么。例如对 Set an alarm for 10 pm意图是set_alarm。本质上是一个**文本分类text classification**问题。槽位填充Slot Filling找出话语中需要被填写的关键信息片段。例如上面的例子中10 pm是触发该意图所需的时间槽位。本质上是一个**序列标注sequence labeling**问题。两个任务可以分别训练两个独立的模型但大量实践证明联合训练joint model效果更好因为意图和槽位在语义上相互约束例如 flight 意图下 colorado 更可能是fromloc.city_name而非meal共享文本表示层可以互相增强。PyText 的IntentSlotModel正是为此设计的——它让文档分类doc classification与词级标注word tagging共享嵌入层与文本表示层正如 pytext/models/joint_model.py 类注释所述A joint intent-slot model. This is framed as a model to do document classification model and word tagging tasks where the embedding and text representation layers are shared for both tasks.本教程使用ATIS 数据集。ATIS 是经典的航空旅行信息系统语料包含航班查询类话语及其意图与槽位标注常用于验证 NLU 联合模型。2. 数据准备将 ATIS 原始语料转换为 PyText 的 TSV 格式2.1 PyText 数据处理器期望的格式PyText 内置的数据处理器data-handler期望数据以tab 分隔的文本文件存储每行包含三个核心字段意图标签intent label、槽位标签slot label和原始话语raw utterance。本教程中TSVDataSource的field_names进一步扩展为五个字段label、slots、text、doc_weight、word_weight后两个是可选的样本权重未提供时模型会使用默认损失权重见下文源码分析。2.2 下载并预处理 ATIS从 Kaggle 下载 ATIS 数据集atis.zip需要免费 Kaggle 账号后用以下两步完成预处理$ unzip download_dir/atis.zip -d download_dir/atis $ python3 demo/atis_joint_model/data_processor.py \ --download-folder download_dir/atis --output-directory demo/atis_joint_model/脚本 demo/atis_joint_model/data_processor.py 会读取下载目录中的四份原始文件输入文件内容atis.dict.intent.csv意图标签词典每行一个行号即 idatis.dict.slots.csv槽位标签词典IOB 格式如B-fromloc.city_name/I-fromloc.city_nameatis.dict.vocab.csv词表atis.train.intent.csv/atis.train.slots.csv/atis.train.query.csv训练集三份数值化文件意图 / 槽位 / 查询词序列atis.test.intent.csv/atis.test.slots.csv/atis.test.query.csv测试集三份数值化文件处理逻辑要点可从源码确认read_vocab()把词典文件逐行读成{行号: 词}的映射用于把数值化的句子还原为真实文本stringify()按空格拆分数值化序列并同时记录每个词在还原句子中的[起始, 结束)字符偏移char offsetsis_valid_slot()/extract_slot_name()处理 IOB 槽位标记只接受B-和I-开头并抽取槽位名get_all_slots()把连续的槽位片段聚合成起始偏移:结束偏移:槽位名的形式多个槽位用逗号连接process_train_set()按VALIDATION_SPLIT 0.25的随机比例把训练数据随机切分成训练集与验证集代码第 11 行常量第 123 行random.random() VALIDATION_SPLIT决定落盘到哪个文件最终写出三个文件atis.processed.train.csv、atis.processed.val.csv、atis.processed.test.csv。处理完成后的每行形如flightTAB0:6:fromloc.city_name,8:11:destloc.city_nameTABflights from colorado to denver脚本还提供-v/--verbose参数开启后会打印每个文件的前 5 行SAMPLE_PRINT_COUNT 5示例方便核对数据是否正确。如果你有自定义数据格式也可以为它编写自定义>$ curl https://nlp.stanford.edu/data/wordvecs/glove.6B.zip demo/atis_joint_model/glove.6B.zip $ unzip demo/atis_joint_model/glove.6B.zip -d demo/atis_joint_model下载的glove.6B.zip压缩包约800 MB。解压后我们会用到其中的glove.6B.100d.txt——即 100 维词向量文件与配置中word_embedding.embed_dim: 100和pretrained_embeddings_path对应。4. 模型配置详解从零训练一个联合 Intent-Slot 模型训练 PyText 模型需要选择正确的任务task与模型架构model architecture。PyText 为大量参数提供了默认值多数情况下能给出合理结果。下面是教程给出的基础版配置可以直接保存为 JSON 文件用于训练{ config: { task: { IntentSlotTask: { data: { Data: { source: { TSVDataSource: { field_names: [ label, slots, text, doc_weight, word_weight ], train_filename: demo/atis_joint_model/atis.processed.train.csv, eval_filename: demo/atis_joint_model/atis.processed.val.csv, test_filename: demo/atis_joint_model/atis.processed.test.csv } }, batcher: { PoolingBatcher: { train_batch_size: 128, eval_batch_size: 128, test_batch_size: 128, pool_num_batches: 10000 } }, sort_key: tokens, in_memory: true } }, model: { representation: { BiLSTMDocSlotAttention: { pooling: { SelfAttention: {} } } }, output_layer: { doc_output: { loss: { CrossEntropyLoss: {} } }, word_output: { CRFOutputLayer: {} } }, word_embedding: { embed_dim: 100, pretrained_embeddings_path: demo/atis_joint_model/glove.6B.100d.txt } }, trainer: { epochs: 20, optimizer: { Adam: { lr: 0.001 } } } } } } }4.1 关键参数含义结合仓库源码逐层解读该配置IntentSlotTaskpytext/task/tasks.py 中定义联合训练文档分类与词级标注任务的专用 Task其Config默认装配IntentSlotModel.Config与IntentSlotMetricReporter.Config。任务通过 pytext/builtin_task.py 的register_builtin_tasks()注册为内置任务因此配置顶层写IntentSlotTask即可被pytext命令行识别。TSVDataSource告诉 PyText 从三个 TSV 文件读取训练 / 验证 / 测试数据并按field_names解析每一列。PoolingBatcher批处理配置。三个 batch size 均为 128pool_num_batches: 10000表示预池化的批次池大小PyText 通过对相似长度样本分桶池化来减少 padding 浪费提高训练效率。sort_key: tokensin_memory: true按 token 数排序以优化批次内长度一致性并把数据载入内存以加速迭代。BiLSTMDocSlotAttention表示层representation layer。文档指出这是一个 BiLSTM 模型其pooling属性决定文档级注意力技术。从 pytext/models/representations/bilstm_doc_slot_attention.py 源码可确认内部lstm默认为BiLSTM也支持OrderedNeuronLSTM、AugmentedLSTM顶层dropout默认 0.4pooling支持SelfAttention、MaxPool、MeanPool此处用SelfAttention对 LSTM 各时刻输出做自注意力得到固定长度的文档表示use_doc_attention逻辑见源码 81–94 行还可选配slot_attention词级注意力与doc_mlp_layers/word_mlp_layers投影层数。output_layer输出层对不同子任务使用不同损失函数doc_output用CrossEntropyLoss交叉熵做文档分类word_output用CRFOutputLayer条件随机场做槽位序列标注。源码 pytext/models/output_layers/intent_slot_output_layer.py 中IntentSlotOutputLayer.Config的word_output类型就是WordTaggingOutputLayer.Config或CRFOutputLayer.Config二选一。word_embedding提供embed_dim与pretrained_embeddings_path指定预训练词向量文件。IntentSlotModel的create_embedding()pytext/models/joint_model.py会基于tokenstensorizer 的词表构建WordEmbedding。trainerepochs: 20优化器Adam学习率0.001。4.2 训练命令在已激活 PyText 环境教程使用名为pytext的 conda/虚拟环境后(pytext) $ pytext train sample_config.json5. 超参数调优与最终结果调优超参数hyper-parameters是获得最佳精度的关键。教程指出通过对学习率、LSTM 层数、隐藏维度、dropout等做超参搜索hyper-parameter sweeps可以拿到槽位标签 F1 约 95%的结果接近该任务当时的 state-of-the-art。教程提到的调优后配置位于demos/atis_intent_slot/atis_joint_config.json——注意该路径为原文档笔误仓库中实际文件是demo/atis_joint_model/atis_joint_config.json。完整的调优配置如下{ config: { task: { IntentSlotTask: { model: { representation: { BiLSTMDocSlotAttention: { lstm: { BiLSTM: { dropout: 0.5, lstm_dim: 366, num_layers: 2, bidirectional: true } }, pooling: { SelfAttention: { attn_dimension: 128 } } } }, word_embedding: { embed_dim: 100, pretrained_embeddings_path: demo/atis_joint_model/glove.6B.100d.txt, embedding_init_strategy: zero }, output_layer: { doc_output: { loss: { CrossEntropyLoss: {} } }, word_output: { CRFOutputLayer: {} } } }, trainer: { epochs: 30, optimizer: { Adam: { lr: 0.001, weight_decay: 0 } } }, data: { Data: { source: { TSVDataSource: { field_names: [ label, slots, text, doc_weight, word_weight ], train_filename: demo/atis_joint_model/atis.processed.train.csv, eval_filename: demo/atis_joint_model/atis.processed.val.csv, test_filename: demo/atis_joint_model/atis.processed.test.csv } }, batcher: { PoolingBatcher: { train_batch_size: 128, eval_batch_size: 128, test_batch_size: 128, pool_num_batches: 10000 } }, sort_key: tokens, in_memory: true } } } }, save_snapshot_path: /tmp/atis_joint_model.pt, export_caffe2_path: /tmp/atis_joint_model.c2 } }与基础配置相比调优配置的关键差异LSTM 层BiLSTM显式指定lstm_dim: 366隐藏维度、num_layers: 2双层、bidirectional: true双向与dropout: 0.5。这与 pytext/models/representations/bilstm.py 的BiLSTM.Config字段一一对应。自注意力SelfAttention的attn_dimension从默认 64 调整为128默认值见 pytext/models/representations/pooling.py 中SelfAttention.Config.attn_dimension 64。词向量初始化新增embedding_init_strategy: zero指定加载预训练向量时未命中词的初始化策略。训练轮数epochs从 20 增加到30优化器保持 Adamlr: 0.001显式weight_decay: 0。导出路径顶层新增save_snapshot_pathPyTorch 模型快照/tmp/atis_joint_model.pt与export_caffe2_path导出模型/tmp/atis_joint_model.c2供下一步预测使用。使用调优配置训练(pytext) $ pytext train demo/atis_joint_model/atis_joint_config.json6. 生成预测让模型回答真实话语训练完成后可以用pytext的predict子命令对单条话语做推理。示例输入是flights from colorado(pytext) $ pytext --config-file demo/atis_joint_model/atis_joint_config.json \ predict --exported-model /tmp/atis_joint_model.c2 {text: flights from colorado}模型的响应是不同意图与槽位的对数概率log probabilities正确意图与槽位应当获得最高分。以下节选展示了预测输出省略号处为被截断的其他类别{ .... doc_scores:flight: array([-0.00016726], dtypefloat32), doc_scores:ground_serviceground_fare: array([-25.865768], dtypefloat32), doc_scores:meal: array([-17.864975], dtypefloat32), .., word_scores:airline_name: array([[-12.158762], [-15.142928], [ -8.991585]], dtypefloat32), word_scores:fromloc.city_name: array([[-1.5084317e01], [-1.3880151e01], [-1.4416825e-02]], dtypefloat32), word_scores:fromloc.state_code: array([[-17.824356], [-17.89767 ], [ -9.848984]], dtypefloat32), word_scores:meal: array([[-15.079164], [-17.229427], [-17.529446]], dtypefloat32), word_scores:transport_type: array([[-14.722928], [-16.700478], [-13.4414 ]], dtypefloat32), ... }输出解读doc_scores:*每个意图的对数概率。doc_scores:flight得分-0.00016726显著高于其他意图meal为 -17.86、ground_serviceground_fare为 -25.87因此模型判定意图为flight。word_scores:*每个词在每个槽位上的得分数组长度为话语的 token 数此处为 3。观察word_scores:fromloc.city_name的第三个值-1.4416825e-02≈ -0.0144接近 0 即为高概率明显高于同一槽位前两个词-15.08 / -13.88也高于其他槽位airline_name、meal、transport_type等均为较大的负数——所以第三个词colorado被标注为出发城市槽位fromloc.city_name。这正是联合模型的价值一次推理同时给出意图整句级与每个词的槽位词级预测。7. 源码级原理联合模型的内部构成为了更好地理解上面配置与预测输出的对应关系可以从源码确认以下实现事实模型装配IntentSlotModel.from_config()pytext/models/joint_model.py依次构建EmbeddingList词嵌入→BiLSTMDocSlotAttention表示层传入embed_dim输出doc_representation_dim/word_representation_dim→IntentSlotModelDecoder分别映射到意图标签数与槽位标签数→IntentSlotOutputLayer。配置里可选的default_doc_loss_weight: 0.2与default_word_loss_weight: 0.5是意图与槽位两部分损失的默认权重若数据未提供doc_weight/word_weight列get_weights_context()会按这两个默认值构造权重向量。表示层BiLSTMDocSlotAttentionpytext/models/representations/bilstm_doc_slot_attention.py默认输出为 LSTM 最后一层每个时刻的特征配置pooling后文档表示变为注意力池化结果outputs[0] self.projection_d(self.doc_attention(lstm_output))词表示默认即 LSTM 逐时刻输出。共享的 LSTM 让两个任务获得同一套上下文表示。输出层与损失IntentSlotOutputLayer.get_loss()pytext/models/output_layers/intent_slot_output_layer.py将意图交叉熵损失与槽位损失CRF 或词标注损失分别乘以对应权重后相加返回联合损失get_pred()同时返回(doc_pred, word_pred)与(doc_score, word_score)。导出命名IntentSlotModel.get_export_output_names()返回[doc_scores, word_scores]这正是预测输出中两组键名的来源预测阶段若有token_indices上下文query_word_reprs()会把词级 logits 对齐回原始 token 位置。8. 结语与延伸本教程走完了 PyText 联合意图-槽位建模的完整闭环ATIS 数据预处理 → GloVe 预训练词向量 → 基础配置训练 → 超参调优F1 ≈ 95%→ 单条话语预测与输出解读。同时通过 demo/atis_joint_model/data_processor.py、demo/atis_joint_model/atis_joint_config.json、pytext/models/joint_model.py 等仓库文件你可以进一步研究 IOB 槽位格式、BiLSTM SelfAttention 表示层、CRF 槽位解码等细节。如果你想深入更多 PyText 建模能力可以继续阅读仓库中的 train_your_first_model.rst、execute_your_first_model.rst 等教程联合模型的指标评估实现可参考 pytext/metric_reporters/intent_slot_detection_metric_reporter.py。请注意本教程对应的旧版 API 已在官方文档中标注过时新项目建议参照仓库中pytext/task/tasks.py的IntentSlotTask与pytext/models/joint_model.py的IntentSlotModel最新配置来搭建实验。赞分享NLP深度学习【免费下载链接】pytextA natural language modeling framework based on PyTorch项目地址https://gitcode.com/gh_mirrors/py/pytext点击查看免费下载相关推荐PyText联合模型完整指南如何实现意图识别与槽位填充的完美结合PyText联合模型完整指南如何实现意图识别与槽位填充的完美结合 PyText联合模型是自然语言理解领域的核心技术能够同时处理意图识别和槽位填充两个关键任务NLP深度学习NLP-progress 意图检测与槽位填充Intent Detection and Slot Filling任务基准与 SOTA 模型全览NLP progress 意图检测与槽位填充Intent Detection and Slot Filling任务基准与 SOTA 模型全览 意图检测与槽位NLP知识库文档PyText 自定义 DataSource 实战为 ATIS 数据实现专属数据源组件并训练意图分类模型PyText 自定义 DataSource 实战为 ATIS 数据实现专属数据源组件并训练意图分类模型 导读 PyText 默认通过 TSVDataSourcNLP深度学习上一篇terraform-provider-aws使用 aws_elasticache_apply_service_update 将 ElastiCache 服务更新应用到集群与复制组下一篇Adobe-GenP 3.0基于AutoIt的Adobe CC授权验证绕过技术实现创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
阅读完成 · 觉得有帮助?