In this episode, Akshat Bubna, CTO of Modal, discusses the evolution of AI infrastructure tailored for agent experience, highlighting Modal's journey from a runtime platform to a specialized cloud for AI workloads. They explore challenges i…
Firehose
Filtered to tagged “LLM inference” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
artificial intelligence 87continual learning 32AI 24reinforcement learning 14agentic coding 13AI safety 13open-weight models 13AI agents 10existential risk 9AI ethics 8cybersecurity 8ethics 7language models 7machine learning 7natural language processing 6open-source 6Reinforcement learning 6security 6artificial general intelligence 5Diffusion models 5recursive self-improvement 5robotics 5software development 5Agentic AI 4large language models 4mathematics 4multi-agent systems 4Recursive self-improvement 4agentic AI 3agents 3
Reiner Pope delivers a blackboard lecture on the mathematical and hardware principles behind training and serving large language models. He explains how batch size, sparsity, and various parallelism strategies (expert, pipeline) impact late…
LLM trainingLLM inferenceBatch sizeLatency optimizationCost analysisRoofline analysisMemory bandwidthCompute performanceKV CacheSparsityMixture of ExpertsExpert parallelismData center architectureScale up networkScale out networkPipeline parallelismMicrobatchingMemory capacityChinchilla scalingRL generationAPI pricingContext lengthCryptographic ciphersNeural network architectureReversible networks