error handling - Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

Question

Welcome To Ask or Share your Answers For Others

error handling - Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

posted Oct 17, 2021 in Technique[技术] by 深蓝 (71.8m points)

error handling - Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

I'm new to NLTK. I'm getting this error and I've searched around for encoding/decoding and specifically the UnicodeDecodeError but this error seems specific to the NLTK source code.

Here's the error:

Traceback (most recent call last):
  File "A:PythonProjectsTestmain.py", line 2, in <module>
    print(pos_tag(word_tokenize("John's big idea isn't all that bad.")))
  File "A:PythonPythonlibsite-packages
ltkag\__init__.py", line 100, in pos_tag
    tagger = load(_POS_TAGGER)
  File "A:PythonPythonlibsite-packages
ltkdata.py", line 779, in load
    resource_val = pickle.load(opened_resource)
UnicodeDecodeError: 'ascii' codec can't decode byte 0xcb in position 0: ordinal not in range(128)

How do I go around fixing this error?

Here's what causes the error:

from nltk import pos_tag, word_tokenize
print(pos_tag(word_tokenize("John's big idea isn't all that bad.")))

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Reply

深蓝 · Answer 1 · 2021-10-17T03:06:52+0000

replyed Oct 17, 2021 by 深蓝 (71.8m points)

try this... NLTK 3.0.1 with Python 2.7.x

import io
f = io.open(txtFile, 'rU', encoding='utf-8')

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

Categories

error handling - Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

error handling - Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

Please log in or register to add a comment.

Please log in or register to reply this article.

1 Reply

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags