web采集的数据为 %u6B63%u5F0F%u4EBA%u5458,需要读取并转换为python对象,想了下不调用Javascript去eval,只能自己翻译了。
核心代码:
import re
import codecs
pattern = re.compile('%u[0-9A-Z]{4}')
n = codecs.open('d:\new.txt', 'w', 'utf-8')
with open('d:\p', 'r') as f:
for l in f:
for i in pattern.findall(l):
l = l.replace(i, unichr(int(i[2:], 16)))
n.write(l)
n.close()