Large language models: Section catalog
Jump to navigation
Jump to search
A catalog of articles from the Systems Analysis Wiki on the topic of Large Language Models (large language model, LLM).
Website: systems-analysis.ru
Key Articles
- Large language model
- Transformer architecture
- Prompt (language models)
- Prompt engineering
- LLM benchmarks
- LLM evaluation
- LLM quality metrics
Large language models (LLM)
- Large language model
- Theoretical foundations of LLM
- Large language model architectures
- Transformer architecture
- Encoder
- Decoder
- Encoder-only
- Decoder-only
- Encoder-Decoder
- Tokenization
- Token
- Embedding
- Context window
- Training large language models
- Pre-training
- Fine-tuning
- In-Context Learning
- Top-p
- Top-k
- Temperature (LLM)
- LLM hallucinations
- Data distortion and bias
- Contextual forgetting
- Generation bias
- Mixture-of-Experts (MoE)
- Reducing LLM errors
- LLM cost optimization
- Open-weight and closed-weight models
- Constitutional AI
- Explainable AI
- RLHF
- Direct Preference Optimization
- Low-Rank Adaptation (LoRA)
- PEFT
- Vector databases
- Multimodal large language models
- Jailbreaks
- FlashAttention
- FlashAttention-2
- FlashAttention-3
- Stop sequences
- Synthetic data generation
- Multimodal reasoning
- Stochastic parrot
Catalog of large language models (LLM)
- T5 (Google)
- LaMDA (Google)
- PaLM
- BERT (Google)
- Gopher (Google)
- Chinchilla (Google)
- Huawei PanGu
- IBM Granite
- BLOOM
- Mixtral (Mistral AI)
- DBRX (Databricks)
- GPT (OpenAI)
- Claude (Anthropic)
- Gemma (Google)
- Gemini (Google)
- LLaMA (Meta)
- Muse Spark (Meta)
- Mistral (Mistral AI)
- DeepSeek
- Grok (xAI)
- Qwen (Alibaba)
- Phi (Microsoft)
- MAI (Microsoft)
- MiniMax
- Jais (UAE)
- Jamba (AI21 Labs)
- Cohere
- Falcon (UAE)
- Kimi (Moonshot AI)
- ERNIE (Baidu)
- GLM (Zhipu AI)
- Nemotron (NVIDIA)
- MiMo (Xiaomi)
- Seed (ByteDance)
- Nova (Amazon)
- Hunyuan (Tencent)
- Yi (01.AI)
- Hugging Face
- OpenAI's large language models
- Google’s large language models
- Anthropic's large language models
- Large language models: Catalog
Prompt engineering (LLM)
- Prompt
- Prompt engineering
- Prompt and context
- Core prompt engineering techniques
- Retrieval‑Augmented Generation (RAG)
- Chain-of-Thought Prompting
- Few-shot and Zero-shot
- Role Prompting
- Tree of Thoughts (ToT)
- Self-refine prompting
- Self-consistency prompting
- Meta Prompting
- Multi-agent prompting
- Prompt compression
- Program of Thoughts Prompting
- Generated Knowledge Prompting
- Multimodal CoT Prompting
- Graph-of-Thoughts
- Chain-of-Verification
- Toolformer
- Least-to-most Prompting
- Automatic Prompt Engineer (APE)
- ReAct Prompting
- Function Calling
- RAG patterns
- GraphRAG
- MM-RAG (Multimodal RAG)
- Hypothetical Document Embeddings (HyDE)
- Hybrid Retrieval
- Packaging & Context Handling
- Prompt engineering: Section catalog
AI agents (LLM)
Evaluation and metric comparison (LLM)
Benchmarks and datasets (LLM)
- LLM benchmarks
- MMLU Benchmark
- MMLU-Pro Benchmark
- MMMU Benchmark
- MMMU-Pro Benchmark
- HellaSwag Benchmark
- HumanEval Benchmark
- TruthfulQA Benchmark
- MT-Bench benchmark
- GLUE Benchmark
- SuperGLUE
- Humanity's Last Exam
- GPQA Diamond Benchmark
- SciCode Benchmark
- Terminal-Bench
- GSM8K (Grade School Math 8K)
- WinoGrande Benchmark
- IFEval Benchmark
- LiveCodeBench
- ARC-AGI-2
- OSWorld
- τ-bench
- HELM Benchmark
- FrontierMath
- AgentHarm
- SafetyBench
- SWE-bench
- SWE-bench Pro
- SWE-bench Verified
- BIG-bench
- Big-Bench Hard
- MATH Benchmark
- MATH-500
- FLORES-200
- RealToxicityPrompts
- PromptRobust
- BOLD
- BBQ
- LMArena
- Elo ranking of language models