Technical field notes from Fanout WAL vs redo log vs undo log, mappedSystem designWeighted jump consistent hash: what worksSystem designGoodput vs SLO attainment in LLM servingInference engineeringOnline vs offline EPLB for MoE servingInference engineeringLabel smoothing minimum loss, calculatedML mathematicsRouter z-loss for stable MoE trainingAI researchSelf-hosted LLM vs API break-even mathInference engineering · ProW4A8 quantization in QServeInference engineeringJump consistent hash vs hash ringSystem designWrite-ahead log vs transaction logSystem designFP8 attention vs FP8 KV cacheInference engineeringvAttention vs PagedAttention for KV cacheInference engineeringExpert parallelism load balancing with EPLBInference engineeringGoodput vs throughput in LLM inferenceInference engineeringToken choice vs expert choice routing in MoEAI researchWhat is a good cross entropy loss valueML mathematicsAI inference engineering, explained with numbersInference engineeringFlashAttention vs PagedAttention: what each fixesInference engineeringW8A8 vs W4A16 quantization: how to chooseInference engineeringExpert parallelism vs tensor parallelism for MoEInference engineeringWhen KV cache quantization slows inferenceInference engineeringDisaggregated prefill and decode with numbersInference engineering · ProWrite-ahead log explained through one crashSystem designConsistent hashing explained simply with 10 keysSystem design · ProSoftmax temperature explained with numbersML mathematicsFP8 vs INT8 vs AWQ vs GPTQInference engineeringMixture of experts routing explainedAI researchKV cache quantization: memory savings by bitsInference engineeringWhy is FlashAttention faster?Inference engineering · ProHow Grok 4.6 happened: Cursor's data flywheelAI researchTensor parallelism vs pipeline parallelismInference engineering · ProWhat is chunked prefill?Inference engineeringArgmax vs max: choice and value explainedML mathematicsHow to read equations in AI research papersML mathematicsKV-cache memory formula for LLM inferenceInference engineeringSoftmax and cross-entropy from logitsML mathematicsThe Roofline model for AI inferenceInference engineeringThe scaled dot-product attention equationML mathematicsWhat the vertical bar means in mathematicsML mathematicsA robotics roadmap for software engineersSystem design · ProHow to estimate LLM API costsInference engineeringHow to use AI to study for examsAI researchHow to use AI to study without losing the workAI researchLLM inference interview questions that matterInference engineering · ProOpenAI vs Anthropic API pricingInference engineeringSpeculative decoding for faster LLM inferenceInference engineeringThe LLM inference engineer roadmapInference engineeringThe open-source robotics stack, explainedSystem designVision-language-action models for roboticsAI research100 papers to understand software and computingSystem designBuild an AI homelab that works like one computerSystem designContinuous batching for LLM inferenceInference engineering · ProPagedAttention: how vLLM manages the KV cacheInference engineering · ProPrefill vs decode in LLM inferenceInference engineering · ProHow AI memory works: five systems, not oneAI research · ProKimi K3 architecture: KDA, MLA and MoEInference engineering · ProKV cache formula for LLM inference memoryInference engineeringDaily paper summaries for AI engineersAI researchDeep learning for computer vision engineersAI researchMachine learning roadmap for backend engineersInference engineeringML interview prep for software engineersAI researchSystem design prep for senior engineersSystem designAI papers for system design engineersSystem designML engineer roadmap for career switchersAI researchML math roadmap without a math degreeML mathematicsPaper-reading habit for staff engineersAI researchSystem design labs for distributed systemsSystem designDeep learning roadmap for self-taught engineersAI researchML interview prep for data scientistsAI researchML math for software engineersML mathematicsSystem design prep for backend engineersSystem designSystem design prep for new gradsSystem designSLMs are more powerful than you thinkAI research