Transformer architecture — สถาปัตยกรรม Transformer
สถาปัตยกรรม Transformer — คือสถาปัตยกรรม Neural Network ที่เสนอขึ้นในปี 2017 โดยนักวิจัยของ Google ในบทความ «Attention Is All You Need»[1] สถาปัตยกรรมนี้ได้ปฏิวัติวงการประมวลผลภาษาธรรมชาติ (NLP) และกลายเป็นรากฐานของ Large Language Model (LLM) สมัยใหม่ส่วนใหญ่ เช่น BERT, GPT และ Gemini นวัตกรรมสำคัญของ Transformer คือกลไก self‑attention ซึ่งช่วยให้โมเดลสามารถถ่วงน้ำหนักความสำคัญของส่วนต่าง ๆ ของข้อมูลนำเข้าและประมวลผลลำดับข้อมูลแบบขนานได้ โดยละทิ้งความเป็นแบบเวียนซ้ำที่มีอยู่ใน RNN และ LSTM
บริบททางประวัติศาสตร์และพื้นหลัง
ก่อนปี 2017 สถาปัตยกรรมหลักที่ใช้ประมวลผลข้อมูลลำดับ เช่น ข้อความ ได้แก่ Recurrent Neural Network (RNN) และตัวแปรที่พัฒนาแล้ว คือ Long Short-Term Memory (LSTM)
ปัญหาของ RNN/LSTM ที่ Transformer แก้ไข
- ข้อจำกัดของการประมวลผลแบบลำดับ: RNN และ LSTM ประมวลผลข้อมูลทีละ token ซึ่งทำให้ไม่สามารถประมวลผลแบบขนานภายในลำดับเดียวกันได้ และทำให้การฝึกบนข้อมูลปริมาณมากช้าลง
- ปัญหา gradient สูญหายและระเบิด: ในลำดับที่ยาว gradient ที่แพร่กระจายย้อนกลับผ่านหลายขั้นตอนอาจลดลงจนเป็นศูนย์หรือเพิ่มขึ้นอย่างไม่มีขอบเขต ทำให้การเรียนรู้ความสัมพันธ์ระยะยาวเป็นเรื่องยาก
- ความสัมพันธ์ระยะยาว: ข้อมูลจากต้นลำดับอาจสูญหายไปก่อนถึงปลายลำดับ
แนวทางของ Transformer คือ การละทิ้งความเป็นแบบเวียนซ้ำโดยสิ้นเชิง เพื่อหันมาใช้กลไก attention แทน วิธีนี้ให้ ความยาวเส้นทางของความสัมพันธ์แบบคงที่ ระหว่างตำแหน่งใด ๆ () ซึ่งช่วยให้การสร้างแบบจำลองความสัมพันธ์ระยะยาวง่ายขึ้น ในขณะที่การ implement self‑attention พื้นฐานมี ความซับซ้อนในการคำนวณแบบกำลังสอง ตามความยาวลำดับ ()[1][2] ปัญหา gradient สูญหาย/ระเบิดไม่ได้ «หายไป» แต่ บรรเทาลง ด้วย residual connection, LayerNorm และรูปแบบการฝึก โดยในการ implement สมัยใหม่มักใช้ตัวแปร Pre‑LayerNorm (Pre‑LN) เนื่องจากมีความเสถียรในการฝึกมากกว่า[3]
สถาปัตยกรรมและองค์ประกอบสำคัญ
สถาปัตยกรรม Transformer ดั้งเดิมประกอบด้วยสองส่วนหลัก ได้แก่ encoder และ decoder ทั้งสองส่วนเป็น stack ของเลเยอร์ที่เหมือนกัน ( ในบทความต้นฉบับ)[1]
- โครงสร้างเลเยอร์ encoder: (1) Multi-Head Attention (MHA) แบบ self‑attention, (2) FFN แบบ position-wise; แต่ละเลเยอร์ย่อยล้อมรอบด้วย residual connection และ LayerNorm[1]
- โครงสร้างเลเยอร์ decoder: (1) self‑attention แบบ masked (causal mask ป้องกันการเข้าถึงตำแหน่งที่อยู่ข้างหน้า), (2) cross‑attention ต่อผลลัพธ์ของ encoder, (3) FFN — พร้อม residual และ LayerNorm เช่นกัน[1]
กลไก attention และ Self‑Attention
กลไก attention คำนวณผลรวมถ่วงน้ำหนักของเวกเตอร์ Value โดยน้ำหนักถูกกำหนดจากระดับความเข้ากันได้ของ Key กับ Query ใน Transformer ใช้ Scaled Dot‑Product Attention:
โดยที่ คือมิติของ Key/Query การหารด้วย ป้องกันการอิ่มตัวของ softmax[1] เมื่อ , และ มาจากลำดับเดียวกัน กลไกนี้เรียกว่า self‑attention
Multi‑Head Attention (ความสนใจแบบหลายหัว)
แทนที่จะใช้ชุดเมทริกซ์ ชุดเดียว จะใช้ «หัว» แบบขนาน แต่ละหัวจะฉาย ไปยังปริภูมิย่อยที่มีมิติน้อยกว่า คำนวณ attention แยกกัน จากนั้นนำผลลัพธ์มาต่อกัน (concatenate) และฉายออกมา[1]:
ตัวแปรสำหรับเร่งความเร็วการ inference:
- MQA (Multi‑Query Attention): ทุกหัวใช้ key/value ชุดเดียวร่วมกัน → ลดปริมาณและ traffic ของ KV‑cache ระหว่างการ decode อย่างมีนัยสำคัญ[4]
- GQA (Grouped‑Query Attention): เป็นการประนีประนอมระหว่าง MHA และ MQA โดยกลุ่มหัวหลายกลุ่มใช้ K/V ร่วมกัน คุณภาพใกล้เคียง MHA แต่มีความเร็วเทียบเท่า MQA[5]
Positional Encoding (การเข้ารหัสตำแหน่ง)
เนื่องจาก self‑attention ไม่รับรู้ลำดับของ token จึงมีการเพิ่ม positional encoding เข้าไปใน embedding ของข้อมูลนำเข้า
- การเข้ารหัสแบบ sinusoidal ดั้งเดิม (PE) จาก[1]:
- ตัวแปรแบบสัมพัทธ์/แบบหมุนสมัยใหม่:
FFN แบบ position-wise, residual และการ normalize
แต่ละเลเยอร์ของ encoder และ decoder นอกจาก attention แล้ว ยังมี FFN แบบ position-wise:
รอบ ๆ แต่ละเลเยอร์ย่อยจะใช้ residual connection และ Layer Normalization: ในเวอร์ชันดั้งเดิมใช้รูปแบบ Post‑LN[1] ใน LLM สมัยใหม่มักใช้ Pre‑LN เพื่อความเสถียรในการฝึกที่ดีกว่าและลดการพึ่งพา warm‑up ยาว[3]
วิวัฒนาการและตัวแปรสมัยใหม่
วิธีการในยุคแรกที่อิงกับ RNN และตัวแปรขั้นสูงอย่าง LSTM ประมวลผลข้อความแบบลำดับทีละ token แม้แนวทางนี้จะสอดคล้องกับโครงสร้างของภาษาได้โดยธรรมชาติ แต่ก็สร้างข้อจำกัดสำคัญ คือ ทำให้การประมวลผลแบบขนานทำได้ยาก และทำให้การค้นหาความสัมพันธ์ระหว่างองค์ประกอบที่อยู่ห่างกันในข้อความเป็นเรื่องยาก ในปี 2017 กลุ่มนักวิจัยจาก Google นำเสนอบทความชื่อ «Attention Is All You Need» โดยอธิบายสถาปัตยกรรมใหม่ที่ชื่อว่า Transformer โมเดลนี้ละทิ้งการใช้ Recurrent Neural Network อย่างสมบูรณ์เป็นครั้งแรก โดยแทนที่ด้วยกลไก attention นวัตกรรมหลักอยู่ที่กลไก attention ช่วยให้ Transformer ประเมินความสำคัญของแต่ละคำในลำดับนำเข้าสำหรับการสร้างคำที่สอดคล้องในผลลัพธ์ได้ ในขณะเดียวกันโมเดลสามารถประมวลผลคำทั้งหมดพร้อมกันได้ ความสามารถในการประมวลผลแบบขนานนี้ทำให้สามารถฝึกโมเดลขนาดใหญ่กว่าบนข้อมูลปริมาณมหาศาลได้ และนำมาสู่การกำเนิด Large Language Model (LLM) สมัยใหม่
สถาปัตยกรรม Transformer เป็นรากฐานของโมเดลจำนวนมาก ซึ่งสามารถแบ่งออกเป็นสามประเภทหลัก
1. โมเดลแบบ Encoder เท่านั้น (Encoder‑only)
- ตัวอย่าง: BERT (และ RoBERTa, ALBERT)[8]
- หลักการ: ฝึกล่วงหน้าด้วยงาน Masked Language Modeling (MLM) พร้อม context แบบสองทิศทาง
- การประยุกต์ใช้: งานทำความเข้าใจ (การจำแนก, NER เป็นต้น)
2. โมเดลแบบ Decoder เท่านั้น (Decoder‑only)
- ตัวอย่าง: ซีรีส์ GPT (GPT‑1/2/3)[9][10], LLaMA[11], Claude
- หลักการ: Causal Language Modeling (CLM) — การทำนาย token ถัดไป โดย attention ถูกกำหนดด้วย causal mask[1]
- การประยุกต์ใช้: การสร้างข้อความ บทสนทนา โค้ด
3. โมเดลแบบ Encoder‑Decoder
- ตัวอย่าง: Transformer ต้นฉบับ, T5, BART[1][12]
- หลักการ: encoder สร้างการแทนค่าของข้อมูลนำเข้า decoder สร้างผลลัพธ์และใช้ cross‑attention กับ feature ของ encoder[1]
- การประยุกต์ใช้: งาน seq2seq (การแปลภาษา การสรุปความ เป็นต้น)
4. สถาปัตยกรรมแบบมัลติโมดัลและทางเลือกอื่น
- Vision Transformer (ViT) — การปรับใช้กับภาพ (การแบ่งเป็น patch)[13]; Swin Transformer — โมเดลแบบลำดับชั้นที่ใช้ shifted windows[14]
- ทางเลือกสำหรับลำดับยาว:
เทคนิคการฝึกและการปรับแต่ง
ประสิทธิภาพของ Transformer มีความเชื่อมโยงอย่างใกล้ชิดกับเทคนิคการฝึกและโครงสร้างพื้นฐาน
- กลยุทธ์การฝึกล่วงหน้า: CLM และ MLM รวมถึง objective แบบ contrastive และ denoising (ELECTRA, T5)[12]
- เทคนิคการ fine-tuning:
- การ fine-tuning เต็มรูปแบบ ของพารามิเตอร์ทั้งหมด
- Parameter-Efficient Fine-Tuning (PEFT): LoRA เพิ่ม adapter แบบ low-rank โดยน้ำหนักพื้นฐานถูก freeze ไว้[18]
- การปรับพฤติกรรม: RLHF — Reinforcement Learning from Human Feedback[19]
- การปรับแต่งระบบสำหรับการ inference: PagedAttention/vLLM เพิ่ม throughput ของการให้บริการด้วยการจัดการ KV‑cache แบบ paging มีประโยชน์อย่างยิ่งสำหรับลำดับยาวและ batch ขนาดใหญ่[20]
ลิงก์
- The Illustrated Transformer — คำอธิบายสถาปัตยกรรม Transformer แบบภาพ
บรรณานุกรม
- Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need. NeurIPS. arXiv:1706.03762.
- Devlin, J., Chang, M.‑W., Lee, K., Toutanova, K. (2019). BERT: Pre‑training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.
- Radford, A., Narasimhan, K., Salimans, T., Sutskever, I. (2018). Improving Language Understanding by Generative Pre‑Training. OpenAI Technical Report.
- Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language Models Are Few‑Shot Learners. NeurIPS. arXiv:2005.14165.
- Raffel, C., Shazeer, N., Roberts, A., et al. (2019). Exploring the Limits of Transfer Learning with a Unified Text‑to‑Text Transformer. arXiv:1910.10683.
- Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al. (2020). An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale. arXiv:2010.11929.
- Liu, Z., Lin, Y., Cao, Y., et al. (2021). Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. arXiv:2103.14030.
- Tay, Y., Dehghani, M., Bahri, D., Metzler, D. (2020). Efficient Transformers: A Survey. arXiv:2009.06732.
- Xiong, R., Yang, Y., He, D., et al. (2020). On Layer Normalization in the Transformer Architecture. ICML. arXiv:2002.04745.
- Su, J., Lu, Y., Pan, S., et al. (2021). RoFormer: Rotary Position Embedding. arXiv:2104.09864.
- Press, O., Smith, N. A., Lewis, M. (2021). Train Short, Test Long: Attention with Linear Biases (ALiBi). arXiv:2108.12409.
- Shazeer, N. (2019). Fast Transformer Decoding: One Write‑Head is All You Need (MQA). arXiv:1911.02150.
- Ainslie, J., Lee‑Thorp, J., de Jong, M., et al. (2023). GQA: Training Generalized Multi‑Query Transformer Models from Multi‑Head Checkpoints. EMNLP. arXiv:2305.13245.
- Kwon, W., Li, Z., Zhuang, S., et al. (2023). Efficient Memory Management for LLM Serving with PagedAttention (vLLM). arXiv:2309.06180.
- Touvron, H., Lavril, T., Izacard, G., et al. (2023). LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971.
- Gu, A., Dao, T. (2023). Mamba: Linear‑Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752.
- Peng, B., et al. (2023). RWKV: Reinventing RNNs for the Transformer Era. arXiv:2305.13048.
- Lieber, O., Lenz, B., Bata, H., et al. (2024). Jamba: A Hybrid Transformer‑Mamba Language Model. arXiv:2403.19887.
- Hu, E. J., Shen, Y., Wallis, P., et al. (2021). LoRA: Low‑Rank Adaptation of Large Language Models. arXiv:2106.09685.
- Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. OpenReview.
หมายเหตุ
- ↑ 1.00 1.01 1.02 1.03 1.04 1.05 1.06 1.07 1.08 1.09 1.10 1.11 Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention Is All You Need. NeurIPS. arXiv:1706.03762.
- ↑ Tay, Y., Dehghani, M., Bahri, D., Metzler, D. (2020). Efficient Transformers: A Survey. arXiv:2009.06732.
- ↑ 3.0 3.1 Xiong, R., Yang, Y., He, D., et al. (2020). On Layer Normalization in the Transformer Architecture. ICML. arXiv:2002.04745.
- ↑ Shazeer, N. (2019). Fast Transformer Decoding: One Write‑Head is All You Need. arXiv:1911.02150.
- ↑ Ainslie, J., Lee‑Thorp, J., de Jong, M., et al. (2023). GQA: Training Generalized Multi‑Query Transformer Models from Multi‑Head Checkpoints. EMNLP. arXiv:2305.13245.
- ↑ Su, J., Lu, Y., Pan, S., et al. (2021). RoFormer: Rotary Position Embedding. arXiv:2104.09864.
- ↑ Press, O., Smith, N. A., Lewis, M. (2021). Train Short, Test Long: Attention with Linear Biases (ALiBi). arXiv:2108.12409.
- ↑ Devlin, J., Chang, M.‑W., Lee, K., Toutanova, K. (2019). BERT: Pre‑training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.
- ↑ Radford, A., Narasimhan, K., Salimans, T., Sutskever, I. (2018). Improving Language Understanding by Generative Pre‑Training. OpenAI.
- ↑ Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language Models Are Few‑Shot Learners. NeurIPS. arXiv:2005.14165.
- ↑ Touvron, H., Lavril, T., Izacard, G., et al. (2023). LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971.
- ↑ 12.0 12.1 Raffel, C., Shazeer, N., Roberts, A., et al. (2019). Exploring the Limits of Transfer Learning with a Unified Text‑to‑Text Transformer. JMLR. arXiv:1910.10683.
- ↑ Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al. (2020). An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale. ICLR. arXiv:2010.11929.
- ↑ Liu, Z., Lin, Y., Cao, Y., et al. (2021). Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. ICCV. arXiv:2103.14030.
- ↑ Gu, A., Dao, T. (2023). Mamba: Linear‑Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752.
- ↑ Peng, B., et al. (2023). RWKV: Reinventing RNNs for the Transformer Era. arXiv:2305.13048.
- ↑ Lieber, O., Lenz, B., Bata, H., et al. (2024). Jamba: A Hybrid Transformer‑Mamba Language Model. arXiv:2403.19887.
- ↑ Hu, E. J., Shen, Y., Wallis, P., et al. (2021). LoRA: Low‑Rank Adaptation of Large Language Models. arXiv:2106.09685.
- ↑ Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. OpenReview.
- ↑ Kwon, W., Li, Z., Zhuang, S., et al. (2023). Efficient Memory Management for LLM Serving with PagedAttention. arXiv:2309.06180.