텍스트 처리/ RNN [머신러닝/딥러닝]
sample = ['the cat sat on the mat.'] * 문자 for word in sample -> word = t, h, e, c, a, t...... * 단어 for word in sample.split() -> word = the, cat, sat, on...... enumerate : for문 index 같이 나옴 ------------------------------------------------------------------ 1. 원-핫 인코딩 : 희소 벡터(대부분 0으로 채워짐), 고차원(어휘사전에 있는 단어수와 동일) - 단어 수준, 문자 수준, 케라스를 사용한 단어수준(tokenizer), 해싱 기법 2. 단어 임베딩 : 밀집 단어 벡터(희소 벡터와 반대), 더 많은 정..