当前位置：首页 > 编程资源 > 编程问答 >内容正文

编程问答

深入Bert实战(Pytorch)----WordPiece Embeddings

发布时间：2023/12/16 编程问答 55 豆豆

生活随笔收集整理的这篇文章主要介绍了深入Bert实战(Pytorch)----WordPiece Embeddings 小编觉得挺不错的,现在分享给大家,帮大家做个参考.

https://www.bilibili.com/video/BV1K5411t7MD?p=5
https://www.youtube.com/channel/UCoRX98PLOsaN8PtekB9kWrw/videos
深入BERT实战(PyTorch) by ChrisMcCormickAI
这是ChrisMcCormickAI在油管bert，8集系列第二篇WordPiece Embeddings的pytorch的讲解的代码，在油管视频下有下载地址，如果不能翻墙的可以留下邮箱我全部看完整理后发给你。

文章目录

- 加载模型
- 查看Bert中的词汇
- 单字符
- Subwords vs. Whole-words
- 开始子词和中间子词
- 对于姓名来说
- 对于数字来说

加载模型

安装huggingface实现

!pip install pytorch-pretrained-bertimport torch from pytorch_pretrained_bert import BertTokenizer# 加载预训练模型 tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')

查看Bert中的词汇

检索整个“tokens”列表，并将它们写入文本文件，可以仔细阅读它们。

with open("vocabulary.txt", 'w', encoding='utf-8') as f:# For each token... 得到每个单词for token in tokenizer.vocab.keys():# Write it out and escape any unicode characters. # 写并且转义为Unicode字符 f.write(token + '\n')

打印出来可以得到
前999个词序是保留位置，它们的形式类似于[unused957]
1 - [PAD] 截断
101 - [UNK] 未知字符
102 - [CLS] 句首，表示分类任务
103 - [SEP] BERT中分隔两个输入的句子
104 - [MASK] MASK机制
第1000-1996行似乎是单个字符的转储。
它们似乎没有按频率排序(例如，字母表中的字母都是按顺序排列的)。
第一个单词是1997位置的The
从这里开始，这些词似乎是按频率排序的。
前18个单词是完整的单词，第2016位是##s，可能是最常见的子单词。
最后一个完整单词是29612位的 “necessitated”

单字符

下面的代码打印出词汇表中的所有单字符的token，以及前面带’##'的所有单字符的token。

结果发现这些都是匹配集——每个独立角色都有一个“##”版本。有997个单字符标记。

下面的单元格遍历词汇表，取出所有单个字符标记。

one_chars = [] one_chars_hashes = []# For each token in the vocabulary... 遍历所有单字符 for token in tokenizer.vocab.keys():# Record any single-character tokens.记录下来if len(token) == 1:one_chars.append(token)# Record single-character tokens preceded by the two hashes. # 记录##单字符 elif len(token) == 3 and token[0:2] == '##':one_chars_hashes.append(token)

打印单字符的

print('Number of single character tokens:', len(one_chars), '\n')# Print all of the single characters, 40 per row.# For every batch of 40 tokens... for i in range(0, len(one_chars), 40):# Limit the end index so we don't go past the end of the list.end = min(i + 40, len(one_chars) + 1)# Print out the tokens, separated by a space.print(' '.join(one_chars[i:end]))

打印##单字符的

print('Number of single character tokens with hashes:', len(one_chars_hashes), '\n')# Print all of the single characters, 40 per row.按每行40打印# Strip the hash marks, since they just clutter the display.去除## tokens = [token.replace('##', '') for token in one_chars_hashes]# For every batch of 40 tokens...每批40 for i in range(0, len(tokens), 40):# Limit the end index so we don't go past the end of the list.限制结束位置end = min(i + 40, len(tokens) + 1)# Print out the tokens, separated by a space.print(' '.join(tokens[i:end])) Number of single character tokens: 997 ! " # $ % & ' ( ) * + , - . / 0 1 2 3 4 5 6 7 8 9 : ; < = > ? @ [ \ ] ^ _ ` a b c d e f g h i j k l m n o p q r s t u v w x y z { | } ~ ¡ ¢ £ ¤ ¥ ¦ § ¨ © ª « ¬ ® ° ± ² ³ ´ µ ¶ · ¹ º » ¼ ½ ¾ ¿ × ß æ ð ÷ ø þ đ ħ ı ł ŋ œ ƒ ɐ ɑ ɒ ɔ ɕ ə ɛ ɡ ɣ ɨ ɪ ɫ ɬ ɯ ɲ ɴ ɹ ɾ ʀ ʁ ʂ ʃ ʉ ʊ ʋ ʌ ʎ ʐ ʑ ʒ ʔ ʰ ʲ ʳ ʷ ʸ ʻ ʼ ʾ ʿ ˈ ː ˡ ˢ ˣ ˤ α β γ δ ε ζ η θ ι κ λ μ ν ξ ο π ρ ς σ τ υ φ χ ψ ω а б в г д е ж з и к л м н о п р с т у ф х ц ч ш щ ъ ы ь э ю я ђ є і ј љ њ ћ ӏ ա բ գ դ ե թ ի լ կ հ մ յ ն ո պ ս վ տ ր ւ ք ־ א ב ג ד ה ו ז ח ט י ך כ ל ם מ ן נ ס ע ף פ ץ צ ק ר ש ת ، ء ا ب ة ت ث ج ح خ د ذ ر ز س ش ص ض ط ظ ع غ ـ ف ق ك ل م ن ه و ى ي ٹ پ چ ک گ ں ھ ہ ی ے अ आ उ ए क ख ग च ज ट ड ण त थ द ध न प ब भ म य र ल व श ष स ह ा ि ी ो । ॥ ং অ আ ই উ এ ও ক খ গ চ ছ জ ট ড ণ ত থ দ ধ ন প ব ভ ম য র ল শ ষ স হ া ি ী ে க ச ட த ந ன ப ம ய ர ல ள வ ா ி ு ே ை ನ ರ ಾ ක ය ර ල ව ා ก ง ต ท น พ ม ย ร ล ว ส อ า เ ་ ། ག ང ད ན པ བ མ འ ར ལ ས မ ა ბ გ დ ე ვ თ ი კ ლ მ ნ ო რ ს ტ უ ᄀ ᄂ ᄃ ᄅ ᄆ ᄇ ᄉ ᄊ ᄋ ᄌ ᄎ ᄏ ᄐ ᄑ ᄒ ᅡ ᅢ ᅥ ᅦ ᅧ ᅩ ᅪ ᅭ ᅮ ᅯ ᅲ ᅳ ᅴ ᅵ ᆨ ᆫ ᆯ ᆷ ᆸ ᆼ ᴬ ᴮ ᴰ ᴵ ᴺ ᵀ ᵃ ᵇ ᵈ ᵉ ᵍ ᵏ ᵐ ᵒ ᵖ ᵗ ᵘ ᵢ ᵣ ᵤ ᵥ ᶜ ᶠ ‐ ‑ ‒ – — ― ‖ ‘ ’ ‚ “ ” „ † ‡ • … ‰ ′ ″ › ‿ ⁄ ⁰ ⁱ ⁴ ⁵ ⁶ ⁷ ⁸ ⁹ ⁺ ⁻ ⁿ ₀ ₁ ₂ ₃ ₄ ₅ ₆ ₇ ₈ ₉ ₊ ₍ ₎ ₐ ₑ ₒ ₓ ₕ ₖ ₗ ₘ ₙ ₚ ₛ ₜ ₤ ₩ € ₱ ₹ ℓ № ℝ ™ ⅓ ⅔ ← ↑ → ↓ ↔ ↦ ⇄ ⇌ ⇒ ∂ ∅ ∆ ∇ ∈ − ∗ ∘ √ ∞ ∧ ∨ ∩ ∪ ≈ ≡ ≤ ≥ ⊂ ⊆ ⊕ ⊗ ⋅ ─ │ ■ ▪ ● ★ ☆ ☉ ♠ ♣ ♥ ♦ ♭ ♯ ⟨ ⟩ ⱼ ⺩⺼⽥、。〈〉《》「」『』〜あいうえおかきくけこさしすせそたちっつてとなにぬねのはひふへほまみむめもやゆよらりるれろをんァアィイウェエオカキクケコサシスセタチッツテトナニノハヒフヘホマミムメモャュョラリルレロワン・ー一三上下不世中主久之也事二五井京人亻仁介代仮伊会佐侍保信健元光八公内出分前劉力加勝北区十千南博原口古史司合吉同名和囗四国國土地坂城堂場士夏外大天太夫奈女子学宀宇安宗定宣宮家宿寺將小尚山岡島崎川州巿帝平年幸广弘張彳後御德心忄志忠愛成我戦戸手扌政文新方日明星春昭智曲書月有朝木本李村東松林森楊樹橋歌止正武比氏民水氵氷永江沢河治法海清漢瀬火版犬王生田男疒発白的皇目相省真石示社神福禾秀秋空立章竹糹美義耳良艹花英華葉藤行街西見訁語谷貝貴車軍辶道郎郡部都里野金鈴镇長門間阝阿陳陽雄青面風食香馬高龍龸 ﬁ ﬂ ！（），－．／：？～

上面两段代码及 ##+单字符和单字符结果一样

# return True print('Are the two sets identical?', set(one_chars) == set(tokens))

Subwords vs. Whole-words

打印一些词汇的统计数据。

import matplotlib.pyplot as plt import seaborn as sns import numpy as npsns.set(style='darkgrid')# Increase the plot size and font size. sns.set(font_scale=1.5) plt.rcParams["figure.figsize"] = (10,5)# Measure the length of every token in the vocab. 加载每个单词 token_lengths = [len(token) for token in tokenizer.vocab.keys()]# Plot the number of tokens of each length. sns.countplot(token_lengths) plt.title('Vocab Token Lengths') plt.xlabel('Token Length') plt.ylabel('# of Tokens')print('Maximum token length:', max(token_lengths))

统计一下，’##'开头的tokens。

num_subwords = 0subword_lengths = []# For each token in the vocabulary... for token in tokenizer.vocab.keys():# If it's a subword...if len(token) >= 2 and token[0:2] == '##':# Tally all subwordsnum_subwords += 1# Measure the sub word length (without the hashes)length = len(token) - 2# Record the lengths. subword_lengths.append(length)

相对于完整词汇表占据的数量

vocab_size = len(tokenizer.vocab.keys())print('Number of subwords: {:,} of {:,}'.format(num_subwords, vocab_size))# Calculate the percentage of words that are '##' subwords. prcnt = float(num_subwords) / vocab_size * 100.0print('%.1f%%' % prcnt) Number of subwords: 5,828 of 30,522 19.1%

作图统计的结果

sns.countplot(subword_lengths) plt.title('Subword Token Lengths (w/o "##")') plt.xlabel('Subword Length') plt.ylabel('# of ## Subwords')

可以自行查看下错误拼写的例子

'misspelled' in tokenizer.vocab # Right 'mispelled' in tokenizer.vocab # Wrong 'government' in tokenizer.vocab # Right 'goverment' in tokenizer.vocab # Wrong 'beginning' in tokenizer.vocab # Right 'begining' in tokenizer.vocab # Wrong 'separate' in tokenizer.vocab # Right 'seperate' in tokenizer.vocab # Wrong

对于缩写来说

"can't" in tokenizer.vocab # False "cant" in tokenizer.vocab # False

开始子词和中间子词

对于单个字符，既有单个字符，也有对应每个字符的“##”版本。子词也是如此吗?

# For each token in the vocabulary... for token in tokenizer.vocab.keys():# If it's a subword...if len(token) >= 2 and token[0:2] == '##':if not token[2:] in tokenizer.vocab:print('Did not find a token for', token[2:])break

可以查看到第一个返回的##ly在词表中，但是ly不在词表中

Did not find a token for ly'##ly' in tokenizer.vocab # True 'ly' in tokenizer.vocab # False

对于姓名来说

下载数据

!pip install wgetimport wget import random print('Beginning file download with wget module')url = 'http://www.gutenberg.org/files/3201/files/NAMES.TXT' wget.download(url, 'first-names.txt')

编码，小写化，输出长度

# Read them in. with open('first-names.txt', 'rb') as f:names_encoded = f.readlines()names = []# Decode the names, convert to lowercase, and strip newlines. for name in names_encoded:try:names.append(name.rstrip().lower().decode('utf-8'))except:continueprint('Number of names: {:,}'.format(len(names))) print('Example:', random.choice(names))

查看有多少个姓名是在BERT的词表中

num_names = 0# For each name in our list... for name in names:# If it's in the vocab...if name in tokenizer.vocab:# Tally it.num_names += 1print('{:,} names in the vocabulary'.format(num_names))

对于数字来说

# Count how many numbers are in the vocabulary. 统计词汇表中有多少数字 count = 0# For each token in the vocabulary... for token in tokenizer.vocab.keys():# Tally if it's a number.if token.isdigit():count += 1# Any numbers >= 10,000?if len(token) > 4:print(token)print('Vocab includes {:,} numbers.'.format(count))

计算一下在1600-2021中有几个数字在

# Count how many dates between 1600 and 2021 are included. count = 0 for i in range(1600, 2021):if str(i) in tokenizer.vocab:count += 1print('Vocab includes {:,} of 421 dates from 1600 - 2021'.format(count))

总结

以上是生活随笔为你收集整理的深入Bert实战(Pytorch)----WordPiece Embeddings的全部内容，希望文章能够帮你解决所遇到的问题。

如果觉得生活随笔网站内容还不错，欢迎将生活随笔推荐给好友。

上一篇： this指针详解
下一篇： java 调用 delphi_【java