← Назад към блога
LLM example
```js const sentences = [ "i like cats", "i like dogs", "i like coding", "you like cats", "you like dogs", "we like coding", "cats like fish", "dogs like meat", "coding is fun", "javascript is fun" ]; const words = [...new Set(sentences.join(" ").split(" "))]; const V = words.length; const id = w => words.indexOf(w); const W = Array.from({ length: V }, () => Array.from({ length: V }, () => Math.random() * 0.2 - 0.1) ); const lr = 0.1; function softmax(x) { const m = Math.max(...x); const e = x.map(v => Math.exp(v - m)); const s = e.reduce((a, b) => a + b); return e.map(v => v / s); } // REAL TRAINING for (let epoch = 0; epoch < 5000; epoch++) { let loss = 0; for (const sentence of sentences) { const ws = sentence.split(" "); for (let i = 0; i < ws.length - 1; i++) { const input = id(ws[i]); const target = id(ws[i + 1]); const p = softmax(W[input]); loss -= Math.log(p[target] + 1e-9); // Cross-entropy gradient p[target] -= 1; for (let j = 0; j < V; j++) W[input][j] -= lr * p[j]; } } if (epoch % 500 === 0) console.log("epoch:", epoch, "loss:", loss.toFixed(3)); } // PREDICT NEXT WORD function predict(word, count = 5) { const probabilities = softmax(W[id(word)]); return probabilities .map((probability, i) => ({ word: words[i], probability })) .sort((a, b) => b.probability - a.probability) .slice(0, count); } console.log(predict("i")); console.log(predict("like")); console.log(predict("coding")); ``` ## Example After training: ```js predict("i"); ``` Might produce: ```text [ { word: "like", probability: 0.99 }, { word: "cats", probability: 0.00 }, { word: "dogs", probability: 0.00 } ] ``` And: ```js predict("like"); ``` Might produce: ```text [ { word: "cats", probability: 0.45 }, { word: "dogs", probability: 0.35 }, { word: "coding", probability: 0.20 } ] ``` ## What the model learns The model learns relationships such as: ```text i → like you → like we → like like → cats like → dogs like → coding cats → like dogs → like coding → is javascript → is is → fun ``` The training process is: ```text input word ↓ neural parameters W ↓ probability for every word ↓ cross-entropy loss ↓ gradient ↓ update W ↓ repeat thousands of times ``` This is genuine gradient-descent training, although the model is extremely small. It is a **word-level neural language model**, not yet a Transformer. A minimal Transformer would add: ```text Tokenization ↓ Embeddings ↓ Self-Attention ↓ Feed-Forward Network ↓ Softmax ↓ Next-token prediction ↓ Backpropagation ```