---
title: "Multimodal large language models — โมเดลภาษาขนาดใหญ่แบบหลายโหมด"
source: "https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94"
wiki: "systems-analysis.info/int"
article: "Multimodal_large_language_models_—_โมเดลภาษาขนาดใหญ่แบบหลายโหมด"
language: "th"
categories:
  - "Category:Large language models"
  - "Category:Machine learning"
  - "Category:Thai"
revision_id: 4765
wiki_created_at: 2026-09-06T23:39:36Z
wiki_modified_at: 2026-09-06T23:39:36Z
downloaded_at: 2026-09-07T23:04:29Z
---

# Multimodal large language models — โมเดลภาษาขนาดใหญ่แบบหลายโหมด

**โมเดลภาษาขนาดใหญ่แบบหลายโหมด** (อังกฤษ: **Multimodal Large Language Models, MLLMs**) คือกลุ่มของโมเดล AI ที่สามารถประมวลผลและสร้างข้อมูลในโหมดต่าง ๆ ได้ ทั้งข้อความ รูปภาพ เสียง และวิดีโอ<sup>[\[1\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-encord_intro-1)</sup> ต่างจากโมเดลภาษาแบบโหมดเดียวที่ทำงานกับข้อความเพียงอย่างเดียว MLLM รวมข้อมูลจากแหล่งต่าง ๆ เพื่อแก้ปัญหาการทำความเข้าใจและการสร้างเนื้อหาที่ซับซ้อน

แนวคิดหลักของ MLLM คือการสร้าง embedding เวกเตอร์เดียวสำหรับโหมดต่าง ๆ ซึ่งช่วยให้โมเดลสร้างความสัมพันธ์เชิงความหมายระหว่าง เช่น รูปภาพกับคำอธิบายข้อความของรูปภาพนั้น<sup>[\[2\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-acm_survey-2)</sup> ความก้าวหน้าสำคัญที่วางรากฐานให้กับ MLLM ยุคปัจจุบันคือการใช้การเรียนรู้แบบ contrastive เพื่อปรับแนวการแทนค่าทางภาพและข้อความในปริภูมิ feature ร่วมกัน ดังที่นำไปใช้ในโมเดล **CLIP**<sup>[\[3\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-radford2021-3)</sup>

## ประวัติการพัฒนา

### ยุคแรก (2013–2020)

รากฐานแนวคิดของ AI แบบหลายโหมดถูกวางในปี 2013 เมื่อนักวิจัยจากมหาวิทยาลัย Stanford แสดงให้เห็นถึงความเป็นไปได้ของการเรียนรู้แบบ zero-shot learning โดยใช้การแทนค่าเวกเตอร์ของคำ<sup>[\[4\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-deoldify2013-4)</sup> ในปี 2016 ทีม **FAIR** (Meta AI) แสดงให้เห็นถึงประสิทธิภาพของการใช้คำอธิบายภาษาธรรมชาติในการฝึกโมเดล computer vision โดยทำความแม่นยำได้ 11.5% บน ImageNet โดยไม่ต้องฝึกตรง ๆ<sup>[\[5\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-openai_fair_2016-5)</sup>

### ยุค CLIP (2021)

จุดเปลี่ยนสำคัญคือการเปิดตัวโมเดล **CLIP** (*Contrastive Language-Image Pre-training*) โดย OpenAI ในเดือนมกราคม 2021 โมเดลที่ฝึกด้วยคู่รูปภาพ-ข้อความ 400 ล้านคู่ แสดงให้เห็นความสามารถในการจำแนกรูปภาพโดยไม่ต้องฝึกเฉพาะทางสำหรับงานใดงานหนึ่ง CLIP กลายเป็นพื้นฐานสำหรับ MLLM ที่ตามมาอีกมากมาย<sup>[\[6\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-stanford_cs_clip-6)</sup>

### การขยายขนาดและนวัตกรรม (2022–2024)

หลังจากความสำเร็จของ CLIP มีโมเดลสำคัญเกิดขึ้นมากมาย:

- **Flamingo** (DeepMind, 2022) — โมเดลขนาด 80 พันล้านพารามิเตอร์ที่แสดงความสามารถโดดเด่นในการเรียนรู้จากตัวอย่างจำนวนน้อย
- **BLIP** (Salesforce, 2022) — สถาปัตยกรรมแบบรวมศูนย์สำหรับการทำความเข้าใจและการสร้างเนื้อหา
- **GPT-4V** (OpenAI, 2023) — โมเดลหลายโหมดเชิงพาณิชย์โมเดลแรกในระดับนี้
- **LLaVA** (Microsoft, 2023) — ทางเลือกแบบโอเพนซอร์สที่ได้รับความนิยมแทน GPT-4V
- **Gemini** (Google, 2023) — สถาปัตยกรรมที่เป็นหลายโหมดโดยกำเนิด ออกแบบมาตั้งแต่ต้นเพื่อรองรับข้อมูลหลายประเภท
- **GPT-4o** (OpenAI, 2024) — โมเดลที่สามารถประมวลผลข้อความ เสียง และวิดีโอแบบ real-time ด้วย latency ต่ำ<sup>[\[1\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-encord_intro-1)</sup>
- **Claude 3.5 Sonnet** (Anthropic, 2024) — โมเดลที่มีความสามารถในการวิเคราะห์ข้อมูลทางภาพที่ดีขึ้น

## แนวทางสถาปัตยกรรม

### สถาปัตยกรรม Dual-Encoder

ใช้ encoder แยกสำหรับแต่ละโหมด ซึ่งฉายข้อมูลเข้าสู่ปริภูมิการแทนค่าร่วม ตัวแทนที่โดดเด่นคือ **CLIP** ซึ่ง visual transformer ประมวลผลรูปภาพและ text transformer ประมวลผลข้อมูลภาษา ข้อดีคือความเป็นโมดูลและประสิทธิภาพในการคำนวณ ข้อเสียคือการปฏิสัมพันธ์ข้ามโหมดที่จำกัด<sup>[\[7\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-viso_ai_mllm-7)</sup>

### สถาปัตยกรรม Encoder-Decoder

Encoder เดียวประมวลผลอินพุตหลายโหมด และ decoder สร้างเอาต์พุตเป็นข้อความ โมเดล **Flamingo** ใช้กลไก *Perceiver Resampler* สำหรับประมวลผลอินพุตทางภาพที่มีความยาวแปรผัน และ layer attention ข้ามโหมด แนวทางนี้ให้การปฏิสัมพันธ์ระหว่างโหมดที่สมบูรณ์ยิ่งขึ้น แต่ต้องการทรัพยากรการคำนวณมากกว่า<sup>[\[8\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-determined_ai_arch-8)</sup>

### สถาปัตยกรรม Alignment

แนวทางนี้ใช้ encoder ที่ผ่านการ pre-training และแช่แข็งไว้ เชื่อมต่อผ่านโมดูล alignment ขนาดเล็กที่ฝึกได้ ตัวอย่างเช่น **BLIP-2** ใช้ **Q-Former** (*Querying Transformer*) เป็นตัวเชื่อมขนาดเบาระหว่าง visual encoder ที่แช่แข็งกับโมเดลภาษา โดยต้องการพารามิเตอร์ที่ฝึกได้น้อยกว่าอย่างมีนัยสำคัญ<sup>[\[9\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-clarifai_blip2-9)</sup>

## โมเดลหลัก

### GPT-4V / GPT-4o (OpenAI)

ตระกูลโมเดล GPT-4 ประเมินว่ามีพารามิเตอร์สูงถึง **1.8 ล้านล้าน** (ในสถาปัตยกรรม mixture of experts) โมเดล **GPT-4o** ที่เปิดตัวในเดือนพฤษภาคม 2024 รองรับการประมวลผลข้อความ รูปภาพ เสียง และวิดีโอแบบ real-time บน benchmark **MMMU** ทำความแม่นยำได้ **69.1%**<sup>[\[10\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-encord_mmmu_perf-10)</sup>

### Gemini (Google)

สถาปัตยกรรมที่เป็นหลายโหมดโดยกำเนิด ฝึกตั้งแต่ต้นด้วยข้อความ รูปภาพ เสียง และวิดีโอ **Gemini 1.5 Pro** รองรับ context window สูงถึง **10 ล้าน token** และเหนือกว่า GPT-4 ใน 30 จาก 32 benchmark ยอดนิยม<sup>[\[11\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-daveai_gemini-11)</sup>

### Claude 3 (Anthropic)

ตระกูลโมเดล (Haiku, Sonnet, Opus) ที่มี context window สูงถึง 200,000 token **Claude 3 Opus** ทำคะแนน **58.5%** บน benchmark MMMU เพื่อเพิ่มความปลอดภัยของโมเดล ใช้แนวทาง Constitutional AI<sup>[\[12\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-anthropic_claude3-12)</sup>

### LLaVA (โมเดลโอเพนซอร์ส)

ผสมผสาน visual encoder CLIP เข้ากับโมเดลภาษา Vicuna มีรุ่นให้เลือกที่มี 7, 13 และ 34 พันล้านพารามิเตอร์ โมเดลทำประสิทธิภาพได้ 85.1% เทียบกับ GPT-4 บนงานสังเคราะห์<sup>[\[13\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-llava_paper-13)</sup>

## ด้านการประยุกต์ใช้

- **การถาม-ตอบด้วยภาพ (VQA)**: ช่วยให้ผู้ใช้สามารถถามคำถามเกี่ยวกับเนื้อหาทางภาพได้
- **การวิเคราะห์เอกสาร**: MLLM ยุคใหม่สามารถประมวลผลได้สูงถึง 2,000 หน้าต่อนาที
- **การสร้างภาพทางการแพทย์**: โมเดลอย่าง **Med-PaLM M** (Google) วิเคราะห์รูปภาพทางการแพทย์และข้อมูลทางคลินิก
- **หุ่นยนต์**: โมเดลอย่าง **RT-2** (Google DeepMind) ช่วยให้หุ่นยนต์เข้าใจสภาพแวดล้อมทางภาพและปฏิบัติตามคำสั่งในภาษาธรรมชาติได้

## ข้อจำกัดในปัจจุบัน

- **Hallucination**: อัตรา hallucination ในเนื้อหาที่สร้างขึ้นประเมินอยู่ที่ 27–46% โมเดลอาจอธิบายวัตถุที่ไม่มีอยู่จริงหรือตีความข้อมูลทางภาพผิดพลาด<sup>[\[14\]](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_note-arxiv_hallucinations-14)</sup>
- **ความต้องการการคำนวณสูง**: การฝึกและการใช้งาน MLLM ต้องการโครงสร้างพื้นฐานด้านการคำนวณที่มีนัยสำคัญ
- **อคติในข้อมูล**: การมีตัวแทนไม่เพียงพอของกลุ่มประชากร ภาษา และวัฒนธรรมในข้อมูลการฝึกนำไปสู่ข้อผิดพลาดเชิงระบบ

## ลิงก์

- A Comprehensive Guide to Multimodal LLMs (Encord Blog)
- Multimodal LLMs: The Complete Guide (Viso.ai)

## เอกสารอ้างอิง

- Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision*. arXiv:2103.00020.
- Alayrac, J.-B. et al. (2022). *Flamingo: a Visual Language Model for Few-Shot Learning*. arXiv:2204.14198.
- Li, J. et al. (2022). *BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation*. arXiv:2201.12086.
- Li, J. et al. (2023). *BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models*. arXiv:2301.12597.
- Liu, H. et al. (2023). *Visual Instruction Tuning*. arXiv:2304.08485.
- Driess, K. et al. (2023). *PaLM-E: An Embodied Multimodal Language Model*. arXiv:2303.03378.
- Brohan, A. et al. (2023). *RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control*. arXiv:2307.15818.
- Yue, X. et al. (2023). *MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI*. arXiv:2311.16502.
- Tsimpoukelli, M. et al. (2021). *Multimodal Few-Shot Learning with Frozen Language Models*. arXiv:2106.13884.
- Singhal, K. et al. (2023). *Med-PaLM 2: Towards Expert-Level Medical Question Answering with Large Language Models*. arXiv:2305.09617.
- Yin, S. et al. (2023). *A Survey on Multimodal Large Language Models*. arXiv:2306.13549.

## หมายเหตุ

1.  <span id="cite_note-encord_intro-1">↑ <sup>[1.0](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-encord_intro_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-encord_intro_1-1)</sup> «A Comprehensive Guide to Multimodal LLMs». *Encord Blog*. <a href="https://encord.com/blog/a-comprehensive-guide-to-multimodal-llms/" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-acm_survey-2">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-acm_survey_2-0) «A Survey on Multimodal Large Language Models». *ACM Computing Surveys*. <a href="https://dl.acm.org/doi/10.1145/3626125" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-radford2021-3">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-radford2021_3-0) Radford, A., et al. «Learning Transferable Visual Models From Natural Language Supervision». *arXiv:2103.00020*. <a href="https://arxiv.org/abs/2103.00020" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-deoldify2013-4">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-deoldify2013_4-0) DeOldify, J. «Zero-Shot Learning by Predicting Attributes». *arXiv:1312.5650*. <a href="https://arxiv.org/abs/1312.5650" class="external autonumber" rel="nofollow">[4]</a></span>
5.  <span id="cite_note-openai_fair_2016-5">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-openai_fair_2016_5-0) «Learning from captions: A milestone in visual language understanding». *OpenAI Blog*. <a href="https://openai.com/blog/learning-from-captions/" class="external autonumber" rel="nofollow">[5]</a></span>
6.  <span id="cite_note-stanford_cs_clip-6">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-stanford_cs_clip_6-0) «Understanding CLIP». *Stanford CS231n*. <a href="https://cs231n.github.io/understanding-visual-representations/" class="external autonumber" rel="nofollow">[6]</a></span>
7.  <span id="cite_note-viso_ai_mllm-7">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-viso_ai_mllm_7-0) «Multimodal LLMs: The Complete Guide». *Viso.ai*. <a href="https://viso.ai/deep-learning/multimodal-llms/" class="external autonumber" rel="nofollow">[7]</a></span>
8.  <span id="cite_note-determined_ai_arch-8">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-determined_ai_arch_8-0) «The Architectures of Multimodal Language Models». *Determined AI*. <a href="https://determined.ai/the-architectures-of-multimodal-language-models/" class="external autonumber" rel="nofollow">[8]</a></span>
9.  <span id="cite_note-clarifai_blip2-9">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-clarifai_blip2_9-0) «Understanding BLIP-2: The New Vision-Language Model». *Clarifai Blog*. <a href="https://www.clarifai.com/blog/understanding-blip-2-the-new-vision-language-model" class="external autonumber" rel="nofollow">[9]</a></span>
10. <span id="cite_note-encord_mmmu_perf-10">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-encord_mmmu_perf_10-0) «MMMU: A New Benchmark for Multimodal LLMs». *Encord Blog*. <a href="https://encord.com/blog/mmmu-benchmark/" class="external autonumber" rel="nofollow">[10]</a></span>
11. <span id="cite_note-daveai_gemini-11">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-daveai_gemini_11-0) «Google Gemini: A Deep Dive». *DaveAI Blog*. <a href="https://dave.ai/blog/google-gemini-a-deep-dive/" class="external autonumber" rel="nofollow">[11]</a></span>
12. <span id="cite_note-anthropic_claude3-12">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-anthropic_claude3_12-0) «Introducing the Claude 3 Family». *Anthropic*. <a href="https://www.anthropic.com/news/claude-3-family" class="external autonumber" rel="nofollow">[12]</a></span>
13. <span id="cite_note-llava_paper-13">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-llava_paper_13-0) Liu, H., et al. «Visual Instruction Tuning». *arXiv:2304.08485*. <a href="https://arxiv.org/abs/2304.08485" class="external autonumber" rel="nofollow">[13]</a></span>
14. <span id="cite_note-arxiv_hallucinations-14">[↑](https://systems-analysis.info/int/Multimodal_large_language_models_%E2%80%94_%E0%B9%82%E0%B8%A1%E0%B9%80%E0%B8%94%E0%B8%A5%E0%B8%A0%E0%B8%B2%E0%B8%A9%E0%B8%B2%E0%B8%82%E0%B8%99%E0%B8%B2%E0%B8%94%E0%B9%83%E0%B8%AB%E0%B8%8D%E0%B9%88%E0%B9%81%E0%B8%9A%E0%B8%9A%E0%B8%AB%E0%B8%A5%E0%B8%B2%E0%B8%A2%E0%B9%82%E0%B8%AB%E0%B8%A1%E0%B8%94#cite_ref-arxiv_hallucinations_14-0) «Hallucinations in Multimodal Large Language Models». *arXiv:2308.08726*. <a href="https://arxiv.org/abs/2308.08726" class="external autonumber" rel="nofollow">[14]</a></span>
