Close Menu
CatchTheBullCatchTheBull
  • Home
  • Crypto News
  • Bitcoin
  • Altcoin
  • Blockchain
  • Airdrops News
  • NFT News
What's Hot

My Decade of Grinding Before Bitcoin

August 16, 2026

Blofin vs MEXC 2026: Choosing Your Best Exchange

August 16, 2026

What Is Flap? Launchpad Paying Stocks to Meme Holders

August 15, 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
CatchTheBullCatchTheBull
  • Home
  • Crypto News
  • Bitcoin
  • Altcoin
  • Blockchain
  • Airdrops News
  • NFT News
CatchTheBullCatchTheBull
Blockchain

Together AI Kernels Team Achieves 3.6x Performance Gains on NVIDIA Hardware

By WebDeskApril 1, 20263 Mins Read
Together AI Kernels Team Achieves 3.6x Performance Gains on NVIDIA Hardware
Share
Facebook Twitter LinkedIn Pinterest Email


Timothy Morano
Apr 01, 2026 19:17

Together AI’s kernel research team delivers major GPU optimization breakthroughs, cutting inference latency from 281ms to 77ms for enterprise AI deployments.





The team behind FlashAttention has quietly become one of the most consequential groups in AI infrastructure. Together AI’s kernel research unit, now about 15 engineers strong, is solving a problem most people don’t even know exists: the massive performance gap between AI models and the hardware running them.

Their latest win? Taking a voice AI company’s time-to-first-token from 281ms down to 77ms—a 3.6x improvement that translated to 7.2x better unit economics.

The Hidden Bottleneck

Here’s what most AI discourse misses: having great models and expensive GPUs doesn’t guarantee performance. The bottleneck sits in between—the kernel layer that translates mathematical operations into actual silicon instructions.

“The gap between what researchers design and what actually runs fast on hardware is vast,” explains Dan Fu, who leads a parallel research lab at UCSD. Get kernels right and you unlock hardware’s full potential. Get them wrong and your expensive GPUs sit partially idle.

For companies building AI-native products, this isn’t academic. When inference costs run 2x higher than necessary, or when latency breaks the user experience, kernel optimization becomes existential.

One Week Versus One Year

The team’s capabilities showed clearly when NVIDIA’s Blackwell GPUs arrived in March 2025. NVIDIA had spent a year with dozens of engineers optimizing kernels for the new architecture. Together AI had a week.

Their secret weapon: ThunderKittens, a library developed with Stanford researchers that reduces kernel code from 1,000+ lines of CUDA to roughly 100-200 lines. The abstraction layer is built around NVIDIA’s tensor cores, the specialized matrix multiplication units on modern GPUs.

Within seven days of hardware access, the team had some of the fastest FP4 and FP8 GEMM kernels available for Blackwell, achieving up to 2x speedups over cuBLAS on H100s.

Real-World Impact

The voice AI case study illustrates what this means in production. The customer had a hard constraint: time-to-first-64-tokens above roughly 100ms breaks conversational flow. Their B200 deployment was hitting 281ms.

Together’s team hand-optimized a “Megakernel” implementation—running an entire model in a single kernel, targeting the HBM bandwidth ceiling of NVIDIA H100s. Results on Llama-3.2-1B: 77ms. On Qwen 2.5 1.5B: 127ms, down from 292ms.

The approach traces back to FlashAttention’s original insight. That Memorial Day 2022 paper proved the AI establishment wrong about attention being fully optimized. By applying database systems principles—data locality, memory hierarchies—to transformer attention, the team achieved 2-3x speedups where previous sparsity methods showed only 10% real gains.

Academic-Industry Pipeline

The team operates through an unusual model. Dan Fu runs his UCSD lab on higher-risk fundamental research. Together AI co-founder Tri Dao is at Princeton. Simran Arora is at Caltech. Ideas get de-risked in academia, then productionized at Together AI. PhD students join the company. Interns work on longer-term research in academic labs.

This produces engineers who bridge theory and production—people who, as Fu puts it, “lose sleep over memory access patterns” and “find beauty in data flow diagrams.”

The work isn’t glamorous. No announcements when a kernel optimization lands. Just faster training times, lower costs, higher throughput. But these margins determine whether AI-native products feel instant or sluggish, whether unit economics work or don’t, whether companies scale to millions of users or plateau at thousands.

For enterprise AI deployments where every millisecond matters—and every percentage point of efficiency translates to significant cost savings—this invisible infrastructure layer may be where the real competitive advantage lies.

Image source: Shutterstock


Credit: Source link

Previous ArticleDeepcoin becomes first CEX to integrate Polymarket ‘event contracts’
Next Article Google Quantum Research Narrows Timeline for Breaking Bitcoin Cryptography

Related Posts

AAVE Price Prediction: Dead Weight Below $92 or a Coiled Spring — The Next 30 Days Are Decisive

August 12, 2026

LDO Price Prediction: $0.25 Is Knocking — One Level Stands Between a Bounce and a Breakdown

August 12, 2026

HBAR Price Prediction: $0.07 Is a Coiled Spring — And the Market Is About to Pick a Direction

August 12, 2026
Add A Comment
Leave A Reply Cancel Reply

Top Posts

My Decade of Grinding Before Bitcoin

August 16, 2026

Blofin vs MEXC 2026: Choosing Your Best Exchange

August 16, 2026

What Is Flap? Launchpad Paying Stocks to Meme Holders

August 15, 2026

Subscribe to Updates

Get the latest Crypto, Blockchain and Airdrop News from us to Catch The Bull.

Advertisement Banner

Welcome to CatchTheBull, your trusted source for the latest Crypto News and Airdrops. We bring you real-time updates, expert insights, and opportunities to stay ahead in the crypto world. Discover trending projects, market analyses, and airdrop details all in one place.

Join us on this journey to navigate the ever-evolving blockchain universe!

Facebook X (Twitter) Instagram YouTube
Top Insights

ASX shareholder seeks court action against former directors

Bitunix vs Bybit 2026: A Comprehensive Comparison Guide

Harmony Exploited in Unauthorized Mint of 4 Billion ONE

Get Informed

Subscribe to Updates

Get the latest Crypto, Blockchain and Airdrop News from us to Catch The Bull.

© 2026 CatchTheBull. All Rights Are Reserved.
  • Contact Us
  • Privacy Policy
  • Terms of Use
  • DMCA

Type above and press Enter to search. Press Esc to cancel.

  • bitcoinBitcoin(BTC)$63,072.000.10%
  • ethereumEthereum(ETH)$1,885.000.10%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$604.94-0.70%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.00-0.10%
  • solanaSolana(SOL)$75.21-0.40%
  • tronTRON(TRX)$0.3314450.00%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-0.50%
  • HyperliquidHyperliquid(HYPE)$57.661.30%
  • dogecoinDogecoin(DOGE)$0.0698830.20%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.0130492.80%
  • leo-tokenLEO Token(LEO)$9.344.60%
  • zcashZcash(ZEC)$490.380.50%
  • moneroMonero(XMR)$408.150.00%
  • chainlinkChainlink(LINK)$9.46-0.80%
  • cardanoCardano(ADA)$0.1783300.50%
  • whitebitWhiteBIT Coin(WBT)$54.670.10%
  • stellarStellar(XLM)$0.157501-0.70%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$203.700.30%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • CantonCanton(CC)$0.095245-3.40%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-0.90%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • litecoinLitecoin(LTC)$44.500.70%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • hedera-hashgraphHedera(HBAR)$0.065197-0.60%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • suiSui(SUI)$0.68-0.50%
  • avalanche-2Avalanche(AVAX)$6.33-2.30%
  • tether-goldTether Gold(XAUT)$4,355.220.00%
  • shiba-inuShiba Inu(SHIB)$0.000004-1.70%
  • crypto-com-chainCronos(CRO)$0.047631-0.60%
  • okbOKB(OKB)$104.70-2.60%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.10%
  • nearNEAR Protocol(NEAR)$1.61-1.50%
  • uniswapUniswap(UNI)$3.302.30%
  • pax-goldPAX Gold(PAXG)$4,371.560.00%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0600695.20%
  • BittensorBittensor(TAO)$197.420.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • AsterAster(ASTER)$0.60-0.30%
  • OndoOndo(ONDO)$0.3312652.20%
  • HTX DAOHTX DAO(HTX)$0.000002-0.60%
  • usddUSDD(USDD)$1.000.00%
  • MemeCoreMemeCore(M)$1.143.10%