cathrynlavery/diagram-design
HTML · ★ 14,210 · 🍴 854 · 📈 4,504 stars today
29 editorial diagram types for Claude Code. Self-contained HTML + SVG. No shadows, no Mermaid-slop.
HTML · ★ 14,210 · 🍴 854 · 📈 4,504 stars today
29 editorial diagram types for Claude Code. Self-contained HTML + SVG. No shadows, no Mermaid-slop.
Python · ★ 6,563 · 🍴 692 · 📈 727 stars today
Graph-Native Infrastructure for Context and Accountable AI Systems
Python · ★ 168,975 · 🍴 20,129 · 📈 383 stars today
Public repository for Agent Skills
Python · ★ 4,917 · 🍴 332 · 📈 768 stars today
14MB foundation model for tiny devices; phones, wearables, smart home, and robots.
Swift · ★ 9,826 · 🍴 662 · 📈 187 stars today
Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. A local Wispr Flow alternative. ⭐ helps a ton :) Windows & iOS waitlist open. Linux soon.
Python · ★ 71,016 · 🍴 6,404 · 📈 354 stars today
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
Rust · ★ 2,567 · 🍴 273 · 📈 1,180 stars today
Macro is a unified workspace for teams: email, chat, docs, tasks, agents, calls, and CRM — @-linked together with shared AI memory.
Python · ★ 12,388 · 🍴 1,667 · 📈 166 stars today
holehe allows you to check if the mail is used on different sites like twitter, instagram and will retrieve information on sites with the forgotten password function.
Python · ★ 20,644 · 🍴 3,313 · 📈 278 stars today
SpiderFoot automates OSINT for threat intelligence and mapping your attack surface.
Rust · ★ 1,181 · 🍴 104 · 📈 408 stars today
Switchyard lets LLM applications route traffic across models and providers while preserving native OpenAI and Anthropic API compatibility - enabling flexible model selection, benchmarking, and cost/performance optimization.
TypeScript · ★ 6,534 · 🍴 598 · 📈 380 stars today
Open-source All in One AI agent workspace. Run any agent — Claude Code, Codex — across your tools (100+ integrations + MCP), apps, browser, and files, with shared memory. Built-in models or BYOK.
★ 45,669 · 🍴 3,295 · 📈 411 stars today
Agent skills for Obsidian. Teach your agent to use Obsidian CLI and open formats including Markdown, Bases, JSON Canvas.
Python · ★ 90,832 · 🍴 7,529 · 📈 204 stars today
Animation engine for explanatory math videos
Shell · ★ 145,159 · 🍴 23,483 · 📈 762 stars today
A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.
Python · ★ 8,896 · 🍴 1,405 · 📈 201 stars today
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
TypeScript · ★ 5,362 · 🍴 574 · 📈 221 stars today
Desktop app to generate 3D models from images using local AI — runs entirely on your GPU
Go · ★ 87,983 · 🍴 10,348 · 📈 473 stars today
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
👍 73
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap betw
👍 1
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on
👍 1
Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessing the reliability of estimation results under occlusion. However, most existing HPE methods output
👍 4
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, managers or shared context. We introduce EvoX Genesis (hereafter, Genesis), which instead makes the softw
👍 96
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study s
👍 6
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporal
👍 5
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark f
👍 176
Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims according to evidence, produce publication-ready figures, and maintain consistency across a long generation process. We present Spark-t
👍 7
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locati
👍 9
Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harness surrounding the model. While supervised fine-tuning and reinforcement learning can adapt agents to
👍 3
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modify databases. Agent safety should be a runtime cont
👍 10
Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource. InSight-doc starts
👍 10
Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either train a smaller multi-vector encoder from scratch or distil only the query side; neither yields a compact single-vector r
👍 12
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiLMMT-46-v0.1 models, we apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models and is
👍 181
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural
👍 4
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in
👍 6
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from
👍 6
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator G_{LM}, built from a positive tensor A_{LM} by elem
👍 12
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safety guardrails, and structural con
👍 10
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-w
👍 11
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher latency, and noisier inputs for final report generation.
👍 25
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against un
👍 2
Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this phenomenon from the perspective of internal adaptation dynamics and discover that neurons in pretrai
👍 2
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit t
👍 5
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pare
👍 72
Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time. As skill libraries grow, a central challenge is to expose the smallest sufficient executable context under a limited context budget. Existing systems
👍 6
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regi
👍 181
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is repeatedly modified and reused across long-horizon workflows. Current safety
👍 10
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement d
👍 2
Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents N, not the cognition of any single agent. We turn a statistical-physics observation
@thejessezhang · 85.7K 粉丝 · 270.4K 阅 · 510 赞 · 30 转
Two-thirds of our deployment work is now done autonomously by our own product (via Duet). This post is about why that is, our philosophy on building a product + deployment motion, and how that's
中文介绍 探讨产品构建与部署自动化哲学,分享 Duet 工具在自动化部署中的应用。
@addyosmani · 408.3K 粉丝 · 175.5K 阅 · 509 赞 · 65 转
For much of human history, we've evaluated code quality via code review: someone reads what you wrote and makes sure it's clean, thoughtful, fast, understandable, and tests well. For agents, that
中文介绍 分析代码审查在人类历史中的作用,探讨代码质量在智能体中的应用。
@nvidia · 2.6M 粉丝 · 59.5K 阅 · 548 赞 · 66 转
The new lightweight open model and routing library delivers greater control over AI, data and workflows across edge devices, PCs, workstations, data centers and the cloud. Source: NVIDIA Blog By Kari
中文介绍 介绍 NVIDIA Nemotron 3.5 和 NeMo Switchyard,提升边缘设备至云端 AI 工作流效率。
@GoogleCloudTech · 1.3M 粉丝 · 56.7K 阅 · 505 赞 · 76 转
Agent Plugins is an open, vendor-neutral standard for packaging Agent Skills and the MCP servers they depend on into one portable folder that any compatible client can load. Google is joining the
中文介绍 介绍 Agent Plugins 标准,Google 加入推动智能体技能的标准化。
@dexhorthy · 30.0K 粉丝 · 42.9K 阅 · 635 赞 · 31 转
tl;dr make your agent converse visually instead of in walls of prose. Lighter and faster than HTML, good enough for most dev-work shaped problems. Coding agents are pretty much unreadable The
中文介绍 提出 /show-me 工具,为编码智能体提供可视化对话界面,提升开发效率。
@mattyp · 44.5K 粉丝 · 33.5K 阅 · 651 赞 · 48 转
I’m always chasing better tools. Notes apps, workflows, optimizations, and now, personal agents. It’s an easy trap to fall into - always building the “custom” thing and sacrificing the work as a
中文介绍 介绍 Grok Bot,分享个人智能体构建与工作优化的心得。
@ericzakariasson · 80.4K 粉丝 · 32.2K 阅 · 719 赞 · 52 转
Grok 4.6 is out! I've used it for a few weeks as my daily driver across the normal mix of coding and knowledge work, and built a few projects with it specifically to push on where it holds up. It's
中文介绍 发布 Grok 4.6 版本,分享作为日常工具的体验和项目应用。
@GoogleAIStudio · 191.7K 粉丝 · 30.7K 阅 · 566 赞 · 60 转
Today, we’re building on the progress of our widely used Flash series by introducing Gemini 3.7 Flash, our most intelligent workhorse model yet for coding and agents. This release comes just three
中文介绍 推出 Gemini 3.7 Flash,强化编码和智能体工作的智能模型。
中文介绍 Hugging Face Blog发布Strands Agents、LeRobot和Hugging Face Storage Buckets一体化解决方案,实现记录、训练和部署数据流。
中文介绍 DeepMind Blog推出Gemini 3.7 Flash,提供更快速的数据处理能力。
The police-tech giant Flock is announcing today that it will change officers’ access to its nationwide network of license plate readers, in an apparent effort to quell a growing backlash and win back contracts lost amid concerns about mass surveillance and police abuse. Several changes aim directly
中文介绍 MIT Tech Review AI报道,警察技术公司Flock为平息对大规模监控和警察滥用担忧的反弹,将调整警察对全国车牌识别网络的访问权限。
Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabilities.
中文介绍 OpenAI发布GPT-5.6构建指南,帮助初创公司构建更快速、成本效益更高的AI代理。
Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output tokens per second.
中文介绍 OpenAI展示Ultrafast模式,GPT-5.6 Sol速度提升至14倍,输出速度达到每秒750个token。
OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its global revenue organization and help businesses realize the full value of AI.
中文介绍 OpenAI任命Dali Rajic为首席营收官,领导全球营收组织。
When we set out to talk to kids about artificial intelligence, we thought we knew what we’d hear. We expected some to tell us they were using it to cheat a little, the way Millennials and Gen Xers opened up CliffsNotes or programmed formulas into their TI-82s, and others to share inspiring ways they
中文介绍 MIT Tech Review AI探讨儿童对人工智能的看法,发现孩子们对AI的应用看法多样。
AI teammate category just had its most significant new entrant yet
中文介绍 Latent Space报道,SpaceXAI推出Grok 4.6和Grok @Bot,为AI团队增加新的成员。
中文介绍 Hugging Face Blog分享从ICML复现2200篇论文的经验。
中文介绍 TLDR AI汇总Claude Chrome Cowork、Grok 4.6和DeepSeek v4-Pro-0813等AI工具的最新动态。
Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potential to transform work. But many organizations find that realizing the desired return on investment (ROI) from AI hinges o
中文介绍 MIT Tech Review AI讨论使用可信数据进行AI代理扩展的挑战。
中文介绍 Hugging Face Blog介绍OlmoEarth embeddings,提供用于下游分析的定制嵌入导出功能。
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
中文介绍 DeepMind Blog推出手语到文本(SL2T)模型,为听障和听力受损用户提供新的手语功能。
中文介绍 Hugging Face Blog介绍LFM2.5-VL-3B,提升边缘视觉能力的解决方案。
Speculative Decoding by any other name would distil as sweet
中文介绍 Latent Space分享如何获取推理轨迹的方法。
PM says he also spoke with the US president about Aukus, which ‘remains full steam ahead’. Follow today’s news live Get our breaking news email, free app or daily news podcast Albanese says gambling inducements are ‘over the top’ Albanese said gambling inducements are “over the top and need to be wo
中文摘要 阿尔班斯表示在与特朗普的“实质性”通话中支持关税豁免,并讨论了“Aukus”协议。
Thousands of sailors on the USS Abraham Lincoln have reportedly faced food shortages and broken plumbing, with some considering jumping overboard.
中文摘要 美国海军陆战队USS Abraham林肯号航母上的数千名水兵面临食物短缺和破损的管道问题,一些水手考虑跳海。
A preliminary report by US investigators says an engine fan blade broke on last month's flight from Greece to Germany.
中文摘要 美国调查人员初步报告称,上月从希腊飞往德国的瑞安航空航班上,发动机风扇叶片损坏导致机窗破裂,乘客头部被吸出。
Emergency services said they were working at the scene of a derailment near Lewes, in East Sussex, on Thursday. Eighteen other people were less badly hurt.
中文摘要 周四,在英国东苏塞克斯郡利文斯附近的火车脱轨事件中,18人受轻伤。
Decrees restricting women’s rights to study, work, travel and act independently now threaten to damage Afghanistan permanently, experts say.
中文摘要 五年来塔利班统治下,阿富汗女性从公共生活中消失,专家警告这可能会永久损害阿富汗。
Supporters of Andrew and Tristan Tate gathered outside a Miami detention centre ahead of their bail hearing.
中文摘要 Tate兄弟在迈阿密的保释听证会前,支持者在外围集会。
U.S. ambassador to Israel Mike Huckabee called settlers besieging a West Bank home of an American-Palestinian family "terrorists."
中文摘要 美国驻以色列大使Mike Huckabee抨击围攻巴勒斯坦裔美国人住宅的定居者为“恐怖分子”。
Israeli drone strikes killed at least 2 people and wounded several others in Gaza on Thursday.
Colombia denies politics played a role after Mexico said its military rescuers were refused entry.
中文摘要 墨西哥表示,哥伦比亚拒绝了其地震救援队入境,但哥伦比亚否认政治介入。
The economy dominates voter concerns in the southern African country.
中文摘要 赞比亚选举中,经济问题是选民关注的焦点。
Trustees condemn Kennedy Center’s move to re-add Trump’s name and shut down for extensive renovations.
中文摘要 肯尼迪艺术中心董事会投票决定恢复特朗普的名字,并将中心关闭两年进行翻修。
At least six homes were set alight by an extensive wildfire in Stourbridge
中文摘要 一场广泛的野火席卷了英国斯托布里奇镇,至少有六座房屋被烧毁。
A powerful explosion tore through a munitions faction south of the Italian capital Rome on Thursday.
中文摘要 意大利罗马以南的一家军火库发生巨大爆炸。
Fidel Castro would have turned 100 on Thursday and his legacy continues to overshadow the island.
中文摘要 古巴庆祝菲德尔·卡斯特罗100岁生日,其遗产继续笼罩着这个岛国。
The spill of Russian crude bound for India compounds the environmental damage from spills in and around the Strait of Hormuz, where tankers and oil facilities have come under attack.
中文摘要 俄罗斯原油在驶往印度的途中泄漏,威胁到阿拉伯海鸟类和海龟栖息地,加剧了霍尔木兹海峡及其周边地区的环境损害。
Asian stocks were poised to extend gains Friday as further evidence of moderating US inflation and a pullback in oil prices reinforced bets that the Federal Reserve will refrain from raising interest rates next month.
中文摘要 亚洲股市预期上涨,因美国通胀降温及油价回落,市场预期美联储下月不会加息。
Some traders are reeling from heavy losses after a brutal correction in South Korea's stock market.
中文摘要 韩国股市剧烈波动,部分投资者一个月内损失14,000美元。
The Securities and Exchange Commission canceled a Friday meeting where the agency was expected to unveil new plans for digital assets as landmark crypto legislation remains stalled in Congress.
中文摘要 美国证券交易委员会推迟加密货币监管会议,标志着行业最新挫折。
Kalshi was told by a judge to stop offering most of its prediction market contracts in Washington state, after regulators said they likely constituted an illegal gambling operation.
中文摘要 Kalshi被法官下令停止在华盛顿州提供大部分预测市场合约,监管机构称其可能构成非法赌博。
Three carriages on the Southern Rail service are flipped on to their side
中文摘要 英国火车脱轨事故造成20人受伤。
It’s tough to measure the economy’s true potential
中文摘要 衡量经济真实潜力为何如此困难。
Gemini Space Station Inc., the digital-asset platform led by the billionaires Tyler and Cameron Winklevoss that went public last year just before the crypto market tumbled from record highs, posted a fourth consecutive quarterly loss since the IPO.
中文摘要 Winklevoss兄弟的Gemini平台发布第四季度连续亏损,但收入增加。
Supplement stacks. Red light therapy. Peptide shots. Many people are biohacking their health to try to live longer, better lives. And now, they’re doing the same for their beloved pets. Anna Edney reports. (Source: Bloomberg)
中文摘要 抗衰老运动也波及猫狗,许多人尝试通过生物黑客技术延长宠物的寿命。
Comprehensive cross-platform coverage of the U.S. market close on Bloomberg Television, Bloomberg Radio, and YouTube with Romaine Bostick, Isabelle Lee, Carol Massar and Tim Stenovec. (Source: Bloomberg)
中文摘要 标普500指数因通胀降温创下历史新高。
Bobby Sharma, Founder and Managing Partner at Bluestone Equity Partners, discussed the recent $12.5 billion valuation of the Los Angeles Lakers, highlighting the extraordinary price and the broader trend of sports becoming institutionalized investment assets. Sharma emphasized that the Lakers repres
中文摘要 前NBA高管:洛杉矶湖人队是体育的特殊资产,其估值达到125亿美元。
Arab Gulf companies are making their debut in Venezuela through an offshore natural gas project to be operated by BP Plc.
中文摘要 阿联酋和卡塔尔公司通过BP领导的天然气项目在委内瑞拉首次亮相。
Tyson Foods Inc. said it would close several additional beef plants, in the latest sign of the US beefpacking industry’s restructuring amid a prolonged US cattle shortage.
中文摘要 泰森食品因美国牛肉短缺持续,将关闭更多牛肉加工厂。
7 回复 · 程序员 节点
6 回复 · 程序员 节点
49 回复 · Apple 节点
96 回复 · Linux 节点
20 回复 · Linux 节点
11 回复 · Apple 节点
17 回复 · Apple 节点
39 回复 · Apple 节点
14 回复 · Apple 节点
6 回复 · Apple 节点
该源今日无内容。