#aibenchmarks kết quả tìm kiếm

aiartgallerie

14 thg 8, 2024

OpenAI launches SWE-bench Verified, a human-validated subset of the popular SWE-bench AI benchmark for evaluating software engineering abilities. GPT-4's score more than doubles! 📈How will this impact AI development in software engineering? #AIBenchmarks #SoftwareEngineering

aiartgallerie's tweet image. OpenAI launches SWE-bench Verified, a human-validated subset of the popular SWE-bench AI benchmark for evaluating software engineering abilities. GPT-4's score more than doubles!

📈How will this impact AI development in software engineering?

#AIBenchmarks #SoftwareEngineering

alby13

@alby13

2 thg 1

"You need to have these very hard tasks which produce undeniable evidence. And that's how the field is making progress today, because we have these hard benchmarks, which represent true progress. And this is why we're able to avoid endless debate." #AIbenchmarks -Ilya Sutskever

alby13's tweet image. "You need to have these very hard tasks which produce undeniable evidence. And that's how the field is making progress today, because we have these hard benchmarks, which represent true progress. And this is why we're able to avoid endless debate." #AIbenchmarks

-Ilya Sutskever

Aspen Digital

@AspenDigital

22 thg 9

It's possible for AI tools to advance the UN's #SDGs, but industry needs to align developers' goals with communities' priorities. This requires better #AIBenchmarks. What are benchmarks, and how can they help? B Cavello explains in this video. Learn more: ow.ly/3VJ650WZ3gV

Suvodeep Mishra

@suvodeepmishra1

23 thg 5

Claude 4 is here—and it's a powerhouse. Outperforms GPT-4 and Gemini 2.5 in reasoning, coding, and long-context tasks. Fast, smart, and ready. #Claude4 #AIbenchmarks

suvodeepmishra1's tweet image. Claude 4 is here—and it's a powerhouse. Outperforms GPT-4 and Gemini 2.5 in reasoning, coding, and long-context tasks. Fast, smart, and ready. #Claude4 #AIbenchmarks

Louis-Paul Baril

@LPBaril

14 thg 11

ERNIE is beating GPT and Gemini on key benchmarks While everyone obsesses over OpenAI and Google, Baidu quietly built something better. And most people have no idea. #ERNIE #BaiduAI #AIBenchmarks #GPTvsERNIE #GeminiComparison

LPBaril's tweet image. ERNIE is beating GPT and Gemini on key benchmarks
While everyone obsesses over OpenAI and Google, Baidu quietly built something better.
And most people have no idea.
#ERNIE #BaiduAI #AIBenchmarks #GPTvsERNIE #GeminiComparison

FutureRealmsTech

@future_realms

11 thg 1

Are there any benchmarks for AI creativity? #AI #AIBenchmarks

Andres Vilariño 🇪🇦

@andresvilarino

6 thg 10

#AIBenchmarks: Why Useless, Personalized Agents Prevail #Benchmarks #AI #ArtificialIntelligence #Tech #technology buff.ly/n3ntJKS

andresvilarino's tweet image. #AIBenchmarks: Why Useless, Personalized Agents Prevail

#Benchmarks #AI #ArtificialIntelligence #Tech #technology

buff.ly/n3ntJKS

StartupHakk

@StartupHakk

13 thg 6

AI Benchmarks RIGGED? Shocking Truth Exposed! (ChatGPT, Claude) #AIBenchmarks #ChatGPT #ClaudeAI #OpenAI #GoogleAI #AIModelEvaluation #ArtificialIntelligence #DataScience #MachineLearning #AIFraud

ITmatters

@ITmatters_in

3 thg 4

Now AI Applications Need to Pass New AI Benchmarks @MLCommons #AIbenchmarks #MLCommons #MLPerf #GenerativeAI #GenAI #MachineLearning #PointPainting #RGAT #LLMs itmatterss.in/new-ai-benchma…

ITmatters_in's tweet image. Now AI Applications Need to Pass New AI Benchmarks

@MLCommons

#AIbenchmarks #MLCommons #MLPerf #GenerativeAI #GenAI #MachineLearning #PointPainting #RGAT #LLMs

itmatterss.in/new-ai-benchma…

Priit @ Amperly AI productivity

@amperlycom

13 thg 12

SWE-bench, launched in Oct 2023, challenges AI coding with 2,294 real-world software engineering problems, raising the bar for AI proficiency. #AIBenchmarks #CodingAI @Stanford

amperlycom's tweet image. SWE-bench, launched in Oct 2023, challenges AI coding with 2,294 real-world software engineering problems, raising the bar for AI proficiency. #AIBenchmarks #CodingAI @Stanford

Priit @ Amperly AI productivity

@amperlycom

3 thg 12, 2024

MMLU evaluates LLM performance across 57 subjects in zero-shot or few-shot scenarios, with GPT-4 and Gemini Ultra achieving top scores. #AIbenchmarks #MMLU @Stanford

amperlycom's tweet image. MMLU evaluates LLM performance across 57 subjects in zero-shot or few-shot scenarios, with GPT-4 and Gemini Ultra achieving top scores. #AIbenchmarks #MMLU @Stanford

SVIC Podcast

@svicpodcast

22 thg 12

Epic AI's Frontier Math: The Toughest Benchmark Yet! #EpicAI #FrontierMath #AIBenchmarks #Mathematics #MachineLearning #DataScience #AIChallenges #CriticalThinking #ProblemSolving #Innovation

Priit @ Amperly AI productivity

@amperlycom

2 thg 12, 2024

HELM (Holistic Evaluation of Language Models) benchmarks LLMs across diverse tasks like reading comprehension, language understanding, and math, with GPT-4 currently leading the leaderboard. #AIbenchmarks @Stanford

amperlycom's tweet image. HELM (Holistic Evaluation of Language Models) benchmarks LLMs across diverse tasks like reading comprehension, language understanding, and math, with GPT-4 currently leading the leaderboard. #AIbenchmarks @Stanford

GoatStack.AI

@GoatstackAI

12 thg 4, 2024

A comprehensive view of existing benchmarks for evaluating AI systems' physical reasoning capabilities. #PhysicalReasoning #AIBenchmarks #GeneralistAgents

GoatstackAI's tweet image. A comprehensive view of existing benchmarks for evaluating AI systems' physical reasoning capabilities. #PhysicalReasoning #AIBenchmarks #GeneralistAgents

Sorab Ghaswalla

@SorabGhaswalla

17 thg 12

If you are about to buy an AI product or service, wait. Listen to this before going ahead. aiforreal.substack.com/p/five-key-thi… #aiproducts #aimetrics #aibenchmarks

WinBuzzer

@WBuzzer

22 thg 9

Scale AI Launches ‘SEAL Showdown’ LLM Leaderboard - Can it Dethrone LMArena? #AI #ScaleAI #AIBenchmarks #LMArena #SEALShowdown #Tech winbuzzer.com/2025/09/22/sca…

WBuzzer's tweet image. Scale AI Launches ‘SEAL Showdown’ LLM Leaderboard - Can it Dethrone LMArena?

#AI #ScaleAI #AIBenchmarks #LMArena #SEALShowdown #Tech

winbuzzer.com/2025/09/22/sca…

Liora R. Herman

@tzionit411

23 thg 7

How Accurate Is #AI at Fixing IaC Security Flaws? 🤔 Eye-opening results: many AI models miss the mark—not from lack of power, but focus. Read the article from our friends at @SymbioticSecAI → symbioticsec.ai/blog/cracking-… #AIBenchmarks #CodeSecurity #DevSecOps #IaC #AppSec

tzionit411's tweet image. How Accurate Is #AI at Fixing IaC Security Flaws? 🤔

Eye-opening results: many AI models miss the mark—not from lack of power, but focus.

Read the article from our friends at @SymbioticSecAI →
symbioticsec.ai/blog/cracking-…

#AIBenchmarks #CodeSecurity #DevSecOps #IaC #AppSec

Paul Triolo

@pstAsiatech

18 giờ

Benchmarkers already show Speciale hitting: 🏅 IMO gold 🏅 CMO gold 🏅 IOI top-tier 🏅 ICPC-world-level coding Read that again: The first open-weights Olympiad-tier reasoning model comes from China, not the US. #AIbenchmarks #MathAI #CodingAI

Demz One🎶🎨🎮🐶🤖demzone.eth|tez @omen_collective

@DemzOneMusic

25 thg 11

🧠 Reasoning results shocked people. Claude Opus 4.5 led in logic, math and multi-step planning, but GPT-5.1 & Gemini 3 Pro were right behind it. The gap is razor-thin and shrinking daily. 🤏🔥 #AIbenchmarks

Aitimess.com

@Aitimess

25 thg 11

1M tokens. Full codebase understanding. PhD-level reasoning. Gemini 3 is built for real work, not demo videos. #Gemini3 #GoogleCloud #AIbenchmarks

Muhammad Azhar

@Azharthegreat

24 thg 11

Grok’s dominance across diverse leaderboards is impressive, especially #1 in Token Usage and Programming Usecase – that indicates not just popularity but real developer trust. Curious how Grok’s agentic telecom edge will push real-world AI automation next? #AIbenchmarks…

Jotium

@Jotiumagent

18 thg 11

🤯 Crushing Benchmarks! Gemini 3.0 Pro significantly outperforms 2.5 Pro on *every* major AI benchmark. It even tops the LMArena Leaderboard with an incredible 1501 Elo score! #AIBenchmarks #GeminiPro

Louis-Paul Baril

@LPBaril

14 thg 11

Wiz Consults

@wizconsults

8 thg 11

China’s Open-Source Triumph in AI: Kimi K2 Thinking Rewrites the Rules digitrendz.blog/?p=81781 #AiBenchmarks #GenerativeAI #KimiK2Thinking #MoonshotAi

wizconsults's tweet card. Meta Description: Moonshot AI’s Kimi K2 Thinking outperforms GPT-5 and Claude Sonnet 4.5 on key benchmarks, proving cost innovation and open-source models from China are rewriting the AI frontier.

China AI Open Source Kimi K2 Thinking

Nguồn: digitrendz.blog

Gary Myers

@AiTripleAce

30 thg 10

Reproducibility validated standard Android A* pathfinding stack; uniform heuristics, 8-directional grid, weighted cost 1.0–1.4. SBOL layer calibrated bias 0.97 ± 0.02 across 1 000 runs, zero-drift at six months. Code held pending IP finalization #SBOL #Grokpedia #AIbenchmarks

HiveForgeAI

@HiveForgeAI

29 thg 10

Anthropic's Claude AI outperformed GPT-4 by 15% on reasoning tests, showcasing advanced multi-agent AI workflows. Hive Forge’s Swarms align perfectly to boost such complex automation across teams. Could this shift how enterprises adopt AI orchestration? #AIbenchmarks #ClaudeAI

CeezerTheGeezer

@CeezerTheGeezer

21 thg 10

Gemini 3.0 is a surprise leader in coding/visual benchmarks, beating Sonnet 4.5. Vision models can now reliably tell time... a huge step for multimodal AI. #GoogleGemini #AIBenchmarks

Không có kết quả nào cho "#aibenchmarks"

shivansh Puri

@shivanshpuri35

16 thg 6

📊 Results: 1. ImageNet (256×256) FID: 3.43 using just 1 step (1-NFE) 2. 50–70% better than past best models 3. Matches big multi-step models at just 2 steps! 🤯 #SOTA #AIbenchmarks

shivanshpuri35's tweet image. 📊 Results:

1. ImageNet (256×256) FID: 3.43 using just 1 step (1-NFE)
2. 50–70% better than past best models
3. Matches big multi-step models at just 2 steps! 🤯

#SOTA #AIbenchmarks

shivansh Puri

@shivanshpuri35

10 thg 6

Did it work? YES. 💥 It beat big models in tests. 🧠 Understands images better 🎨 Generates nicer pictures ✅ Even humans liked its results more than others #AIbenchmarks

shivansh Puri

@shivanshpuri35

17 thg 6

Real results 📊 ShiQ did great in tests: ✔️ It learned faster ✔️ Needed less data ✔️ Worked better for multi-turn conversations #AIbenchmarks #AIperformance

shivanshpuri35's tweet image. Real results 📊
ShiQ did great in tests:
✔️ It learned faster
✔️ Needed less data
✔️ Worked better for multi-turn conversations

#AIbenchmarks #AIperformance

alby13

@alby13

2 thg 1

Andres Vilariño 🇪🇦

@andresvilarino

6 thg 10

#AIBenchmarks: Why Useless, Personalized Agents Prevail #Benchmarks #AI #ArtificialIntelligence #Tech #technology buff.ly/n3ntJKS

Mudde AI 🇺🇬

@MuddeAI

22 thg 1

Performance speaks volumes! 📊 DeepSeek-R1 outperforms competitors across benchmarks like AIME 2024 and MATH, showcasing groundbreaking accuracy. #DeepSeekR1 #AIBenchmarks

MuddeAI's tweet image. Performance speaks volumes! 📊 DeepSeek-R1 outperforms competitors across benchmarks like AIME 2024 and MATH, showcasing groundbreaking accuracy.
#DeepSeekR1 #AIBenchmarks

FutureRealmsTech

@future_realms

11 thg 1

Are there any benchmarks for AI creativity? #AI #AIBenchmarks

Louis-Paul Baril

@LPBaril

14 thg 11

BernardLeong.eth

@bernardleong

4 thg 12, 2023

Performance of LLMs: GPT-4 from @OpenAI is still leading the open-sourced ones. #GenerativeAI #AIBenchmarks

aiartgallerie

@aiartgallerie

14 thg 8, 2024

Suvodeep Mishra

@suvodeepmishra1

23 thg 5

Claude 4 is here—and it's a powerhouse. Outperforms GPT-4 and Gemini 2.5 in reasoning, coding, and long-context tasks. Fast, smart, and ready. #Claude4 #AIbenchmarks

ITmatters

@ITmatters_in

3 thg 4

Now AI Applications Need to Pass New AI Benchmarks @MLCommons #AIbenchmarks #MLCommons #MLPerf #GenerativeAI #GenAI #MachineLearning #PointPainting #RGAT #LLMs itmatterss.in/new-ai-benchma…

AI MATT

@aimattant

21 thg 9, 2024

The numbers are in! GPT-4O takes the lead across the board, but Qwen2.5-72B holds its ground. 93.7 vs 86.1 on MMLU, 97.8 vs 91.5 on GSM8K. The AI race is heating up! #GPT4O #Qwen25 #AIBenchmarks

aimattant's tweet image. The numbers are in! GPT-4O takes the lead across the board, but Qwen2.5-72B holds its ground. 93.7 vs 86.1 on MMLU, 97.8 vs 91.5 on GSM8K. The AI race is heating up! #GPT4O #Qwen25 #AIBenchmarks

Vlad Bogolin

@vladbogo

23 thg 1, 2024

🧵 [5/n] The results are quite promising. After three iterations, the enhanced Llama 2 70B model outperformed others models such as Claude 2 and GPT-4 0613 on the AlpacaEval 2.0 leaderboard. #AIBenchmarks #Performance

vladbogo's tweet image. 🧵 [5/n] The results are quite promising. After three iterations, the enhanced Llama 2 70B model outperformed others models such as Claude 2 and GPT-4 0613 on the AlpacaEval 2.0 leaderboard. #AIBenchmarks #Performance

Priit @ Amperly AI productivity

@amperlycom

13 thg 12

SWE-bench, launched in Oct 2023, challenges AI coding with 2,294 real-world software engineering problems, raising the bar for AI proficiency. #AIBenchmarks #CodingAI @Stanford

Priit @ Amperly AI productivity

@amperlycom

3 thg 12, 2024

MMLU evaluates LLM performance across 57 subjects in zero-shot or few-shot scenarios, with GPT-4 and Gemini Ultra achieving top scores. #AIbenchmarks #MMLU @Stanford

®️Investor

@DEAD_IN_PARIS

21 thg 4, 2024

3/7:Mistral's latest model, Mixtral 8x22B, is said to outperform Meta's Llama 2 70B in math and coding tests. Mixture-of-experts architecture FTW! #MixtralLLM #AIBenchmarks #TechCompetition

DEAD_IN_PARIS's tweet image. 3/7:Mistral's latest model, Mixtral 8x22B, is said to outperform Meta's Llama 2 70B in math and coding tests. Mixture-of-experts architecture FTW! #MixtralLLM #AIBenchmarks #TechCompetition

0xWulf

@hexawulf

21 thg 9

🚀 Impressive leap on the #AI leaderboard: Mistral AI's new Magistral models just jumped up the Artificial Analysis Intelligence Index—punching far above their weight and rivaling models many times larger. Size isn’t everything anymore! #AIbenchmarks #LLMs #MistralAI

hexawulf's tweet image. 🚀 Impressive leap on the #AI leaderboard: Mistral AI's new Magistral models just jumped up the Artificial Analysis Intelligence Index—punching far above their weight and rivaling models many times larger. Size isn’t everything anymore! #AIbenchmarks #LLMs #MistralAI

dylan

@dylangiuffrida

28 thg 5, 2024

3/9 What's impressive is that despite its smaller size, Ph-3 Vision exceeds in benchmarks, scoring highly in MMU, MM Bench, Science QA, and more. 🏅📊 Smaller but mighty! #AIBenchmarks

dylangiuffrida's tweet image. 3/9 What's impressive is that despite its smaller size, Ph-3 Vision exceeds in benchmarks, scoring highly in MMU, MM Bench, Science QA, and more.

🏅📊 Smaller but mighty!

#AIBenchmarks

Something went wrong.

United States Trends

1. Cowboys 72.9K posts
2. LeBron 101K posts
3. Gibbs 19.8K posts
4. #heatedrivalry 22.6K posts
5. Lions 90.1K posts
6. Pickens 14.2K posts
7. scott hunter 4,377 posts
8. Brandon Aubrey 7,259 posts
9. #OnePride 10.4K posts
10. Ferguson 10.8K posts
11. #DALvsDET 6,136 posts
12. Shang Tsung 25.7K posts
13. Eberflus 2,610 posts
14. CeeDee 10.5K posts
15. Paramount 19.2K posts
16. fnaf 2 25.2K posts
17. Warner Bros 18.3K posts
18. Goff 8,644 posts
19. Bland 8,590 posts
20. DJ Reed N/A