Tensorflow 的 seq2seq：tensorflow.python.framework.errors_impl.InvalidArgumentError

发布于 2025-01-10 08:49:41 字数 2363 浏览 6 评论 0原文

我在这里非常密切地关注 Seq2seq 翻译教程 https://www.tensorflow.org/addons/tutorials/networks_seq2seq_nmt#define_the_optimizer_and_the_loss_function 在测试其他数据时。我在实例化编码器时遇到错误，该编码器被定义为

class Encoder(tf.keras.Model):
  def __init__(self, vocab_size, embedding_dim, enc_units, batch_sz):
    super(Encoder, self).__init__()
    self.batch_sz = batch_sz
    self.enc_units = enc_units
    self.embedding = tf.keras.layers.Embedding(vocab_size, embedding_dim)

    ##-------- LSTM layer in Encoder ------- ##
    self.lstm_layer = tf.keras.layers.LSTM(self.enc_units,
                                   return_sequences=True,
                                   return_state=True,
                                   recurrent_initializer='glorot_uniform')

  def call(self, x, hidden):
    x = self.embedding(x)
    output, h, c = self.lstm_layer(x, initial_state = hidden)
    return output, h, c

  def initialize_hidden_state(self):
    return [tf.zeros((self.batch_sz, self.enc_units)), tf.zeros((self.batch_sz, self.enc_units))]

在此处进行测试时它正在下降

# Test Encoder Stack
encoder = Encoder(vocab_size, embedding_dim, units, BATCH_SIZE)

# sample input
sample_hidden = encoder.initialize_hidden_state()
sample_output, sample_h, sample_c = encoder(example_input_batch, sample_hidden)

错误是以下

Traceback (most recent call last):
  File "C:/Users/Seq2seq/Seq2seq-V3.py", line 132, in <module>
    sample_output, sample_h, sample_c = encoder(example_input_batch, sample_hidden)
  File "C:\Users\AppData\Local\Programs\Python\Python39\lib\site-packages\keras\utils\traceback_utils.py", line 67, in error_handler
    raise e.with_traceback(filtered_tb) from None
  File "C:/Users/Seq2seq/Seq2seq-V3.py", line 119, in call
    x = self.embedding(x)
tensorflow.python.framework.errors_impl.InvalidArgumentError: Exception encountered when calling layer "embedding" (type Embedding).

indices[12,148] = 106 is not in [0, 106) [Op:ResourceGather]

Call arguments received:
  • inputs=tf.Tensor(shape=(64, 200), dtype=int64)

TF 2.0

这可能是 TF Addons 中的问题，您对此有一些经验吗？

编辑

教程在单词级别标记化：我在字符级别对文本进行编码，106 是我的 vocab_size （字符数）

原文

I am following quite closely the Seq2seq for translation tutorial here https://www.tensorflow.org/addons/tutorials/networks_seq2seq_nmt#define_the_optimizer_and_the_loss_function while testing on other data. I meet an error when instantiating the Encoder which is defined as

class Encoder(tf.keras.Model):
  def __init__(self, vocab_size, embedding_dim, enc_units, batch_sz):
    super(Encoder, self).__init__()
    self.batch_sz = batch_sz
    self.enc_units = enc_units
    self.embedding = tf.keras.layers.Embedding(vocab_size, embedding_dim)

    ##-------- LSTM layer in Encoder ------- ##
    self.lstm_layer = tf.keras.layers.LSTM(self.enc_units,
                                   return_sequences=True,
                                   return_state=True,
                                   recurrent_initializer='glorot_uniform')

  def call(self, x, hidden):
    x = self.embedding(x)
    output, h, c = self.lstm_layer(x, initial_state = hidden)
    return output, h, c

  def initialize_hidden_state(self):
    return [tf.zeros((self.batch_sz, self.enc_units)), tf.zeros((self.batch_sz, self.enc_units))]

It is falling when testing here

# Test Encoder Stack
encoder = Encoder(vocab_size, embedding_dim, units, BATCH_SIZE)

# sample input
sample_hidden = encoder.initialize_hidden_state()
sample_output, sample_h, sample_c = encoder(example_input_batch, sample_hidden)

The error is the following

Traceback (most recent call last):
  File "C:/Users/Seq2seq/Seq2seq-V3.py", line 132, in <module>
    sample_output, sample_h, sample_c = encoder(example_input_batch, sample_hidden)
  File "C:\Users\AppData\Local\Programs\Python\Python39\lib\site-packages\keras\utils\traceback_utils.py", line 67, in error_handler
    raise e.with_traceback(filtered_tb) from None
  File "C:/Users/Seq2seq/Seq2seq-V3.py", line 119, in call
    x = self.embedding(x)
tensorflow.python.framework.errors_impl.InvalidArgumentError: Exception encountered when calling layer "embedding" (type Embedding).

indices[12,148] = 106 is not in [0, 106) [Op:ResourceGather]

Call arguments received:
  • inputs=tf.Tensor(shape=(64, 200), dtype=int64)

TF 2.0

This might be a problem in TF Addons, would you have some experience with that?

EDIT

the tutorial tokenizes at the word level : I encode the text at the char level and 106 is my vocab_size (number of characters)

分享到QQ

分享到微博

如果你对这篇内容有疑问，欢迎到本站社区发帖提问参与讨论，获取更多帮助，或者扫码二维码加入 Web 技术交流群。

发布评论

需要登录才能够评论，你可以免费注册一个本站的账号。

靑春怀旧 2025-01-17 08:49:42

这已经足够暗示了，事实上

indices[12,148] = 106 is not in [0, 106) [Op:ResourceGather]

我必须确保我的词汇量是 vocab_size = len(vocab)+1 。数据集构建现在正在进行

text = open(FILE_PATH, 'rb').read().decode(encoding='utf-8') 
vocab = sorted(set(text))

# [...]

vocab_size = len(vocab)+1

This is enough of a hint in fact

indices[12,148] = 106 is not in [0, 106) [Op:ResourceGather]

I had to make sure my vocabulary is vocab_size = len(vocab)+1. The dataset construction now goes

text = open(FILE_PATH, 'rb').read().decode(encoding='utf-8') 
vocab = sorted(set(text))

# [...]

vocab_size = len(vocab)+1

回复收藏 0 原文

往日 2025-01-17 08:49:42

词汇量大小应始终为：len(vocab)+1

回复收藏 0 原文

~没有更多了~

关于作者

权谋诡计

暂无简介

文章

26 人气

关注发私信

櫻之舞

文章 0 评论 0

关注

弥枳

文章 0 评论 0

关注

m2429

文章 0 评论 0

关注

寻找一个思念的角度

文章 0 评论 0

关注

野却迷人

文章 0 评论 0

关注

我怀念的。

文章 0 评论 0

友情链接

文江博客

Tensorflow 的 seq2seq：tensorflow.python.framework.errors_impl.InvalidArgumentError

如果你对这篇内容有疑问，欢迎到本站社区发帖提问参与讨论，获取更多帮助，或者扫码二维码加入 Web 技术交流群。

发布评论

评论（2）

关于作者

相关话题

热门标签

推荐作者

櫻之舞

弥枳

m2429

寻找一个思念的角度

野却迷人

我怀念的。

友情链接

Tensorflow 的 seq2seq：tensorflow.python.framework.errors_impl.InvalidArgumentError

如果你对这篇内容有疑问，欢迎到本站社区发帖提问 参与讨论，获取更多帮助，或者扫码二维码加入 Web 技术交流群。

发布评论

评论（2）

关于作者

相关话题

热门标签

推荐作者

櫻之舞

弥枳

m2429

寻找一个思念的角度

野却迷人

我怀念的。

友情链接

如果你对这篇内容有疑问，欢迎到本站社区发帖提问参与讨论，获取更多帮助，或者扫码二维码加入 Web 技术交流群。