AI + Tech Dashboard

Research first. Strong opinions second. Live data throughout.

Latest arXiv papers, your opinionated takes, and benchmark context in one page built for trust and repeat visits.

One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

Pipeline figures in ML papers must be repurposed across many canvases, including paper columns, 16:9 slides, portrait posters, 1:1 social teasers, 9:16 phone previews. Each format imposes a different aspect ratio on the same computational graph, where any silently broken connection misrepresents the method. We formulate aspect-ratio-adaptive flowchart relayout as a distinct task: given a raster flowchart and a target ratio, produce a structurally faithful, hallucination-free, editable layout. Existing methods fail characteristically: image-to-image models stretch blocks and reject extreme ratios, text-to-image agentic systems hallucinate content, and parse-then-render systems mis-route edges. We propose an agentic pipeline factored into Parse, Style, and Layout stages, each pairing a main agent with a critic that combines deterministic constraint checks with VLM visual feedback so connectivity is explicitly checked and prevented from being silently broken. Outputs are draw.io-editable mxGraph XML. On a curated benchmark of 100 flowcharts at five aspect ratios, evaluated by Gemini 3.1 Pro and validated against human judgments, our method reaches 68.6% Content Fidelity versus 11.2-41.4% for prior work. Project page: https://onefigureeverycanvas.vercel.app/

Base Models Can Reason By Taking a Cue From Training Data

In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.

InterMimicGen: Scaling Humanoid Loco-Manipulation through Self-Evolving Motion Imitation

Captured human-object interactions provide rich supervision for humanoid loco-manipulation, but they are sparse, heterogeneous, and not directly executable by robots. We introduce InterMimicGen, a self-evolving motion-imitation framework in which robot motion data and a tracking policy improve each other. First, we consolidate motion-captured human-object interaction datasets and retarget them into humanoid robot references while preserving whole-body coordination and dexterous hand-object relationships. This produces a large and diverse humanoid robot reference collection for dexterous whole-body loco-manipulation. Second, we train a physics-based generalist tracker that executes these references in simulation on a humanoid with dexterous hands, covering a scale and diversity beyond prior humanoid tracking systems for loco-manipulation. Third, we close a data flywheel: each round makes small, task-preserving changes to where an interaction takes place and how the body performs it, fine-tunes the tracker on them, and keeps only the variants whose simulated execution completes the task, which seed the next round. With more iterations, these small edits compound into broader coverage around the sparse original demonstrations while preserving task semantics and motion quality. Experiments show contact-preserving retargeting across robot configurations, broad tracking with a single generalist policy, executable motions that keep growing over augmentation rounds, and transfer to real robots. InterMimicGen provides a unified path from heterogeneous human demonstrations to a continually expanding motion resource for humanoid robot learning.

TranScope: What the Software Hides About LLM Training Data, the Hardware Reveals at Scale, and Accelerators Magnify

Membership is the root privacy primitive in machine learning: to date, no hardware-based out-of-distribution detection on black-box models has been demonstrated against constant-time, static neural networks with masked confidence. This paper performs the first cycle-level examination of how large language models and vision transformers interact with various modern microarchitecture components, including integrated accelerators, as LLMs scale in size and answers the question of whether the data that a model was trained on affects its execution footprint even without any input-dependent branch, dynamic optimization, or early exit and in constant-time models. The results confirm that the answer is yes and identify which modern hardware components, such as TLBs or on-core accelerators, reveal or amplify that effect. The results also answer whether the signal is informative enough to reliably classify the in-/vs/out-of-distribution property of membership. To understand why, we perform a systematic root cause analysis and find that the transformer's tokenization steps, which happen during training, alter the locality of the accesses the model makes to fetch the vocabulary token later during inference and, as a result, change the page table access patterns and TLB in a previously unknown data-dependent way, causing microarchitectural state to vary significantly based on whether or not the input was in the distribution of the transformer training data. Building on the above observation, we introduce TranScope: the first microarchitecture tool for detecting membership information with low cost, no need for a surrogate model, and significantly higher robustness, e.g., 0.6 AUC for PETAL (best previously reported) vs 0.9 AUC (ours). This reintroduces hardware as both an opportunity, e.g., a tool for checking copyright violation for the first time, and a new channel for inferring membership (MIA).

S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation

Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which performs autoregressive diffusion at high noise before switching to parallel diffusion at low noise. The autoregressive phase provides the serial computation needed to coordinate interdependent events and produce valid state transitions while the parallel phase jointly refines the entire video and reduces sampling time relative to fully serial generation. We implement S2PD with two architectures: a pixel-space diffusion transformer trained from scratch and a pretrained video model adapted through LoRA fine-tuning with causal attention. Across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines and generates videos with greater temporal stability and sampling efficiency than other serial methods.

BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

Worst-group accuracy (WGA) evaluates a trained predictor but does not characterize how its frozen backbone behaves when a new head is learned. We introduce BiasFlow, a hook-based toolkit for monitoring class-attribute centroid alignment (IBMI), within-class centroid separation (W-IBMI), and feature-projection sensitivity. IBMI is confounded by class-attribute correlation and is not a measure of causal feature reliance. We pair these diagnostics with BiasFlow Regularization (BFR), a supervised, composable class-conditional centroid-alignment penalty. W-IBMI verifies the quantity BFR optimizes; it is scale dependent and does not independently establish attribute removal. Across the reported small-scale benchmarks, adding BFR improves or preserves mean WGA, with gains up to +26.0 pp on UrbanCars. The principal independent stress test freezes CelebA-Std backbones and trains fresh heads on biased data: BFR+GroupDRO improves WGA from 40.7% to 64.1%, while Male probe accuracy decreases from 92.5% to 72.2%. Attribute information remains recoverable, and cross-task results are mixed. A controlled synthetic-watermark ImageNet experiment additionally improves watermark-shift accuracy by +23.0 pp under matched training. These results support evaluating centroid geometry and resistance to biased head retraining alongside WGA, within the tested protocols.

Learning to Read the Contextual Tokens in Diffusion Transformers

Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading this contextual space through natural-language interrogation. We train a lightweight bottleneck network that maps intermediate contextual tokens into the input space of a frozen Large Language Model (LLM), allowing the LLM to answer questions about the emerging image directly from these hidden representations. Our reader reveals that contextual tokens encode a rich, global representation of the emerging scene: generation-specific semantics, including attributes left underspecified by the prompt, are accessible surprisingly early in denoising, while increasingly fine-grained details become readable over time. Remarkably, this information remains decodable even when the MM-DiT receives an empty prompt, showing that contextual tokens accumulate substantial image-specific information from the evolving visual representation itself. We further find that generations with more readable contextual representations tend to receive higher human-preference scores. Building on these observations, we introduce Contextual Alignment, a training technique that explicitly reinforces the visual-semantic information encoded in the contextual tokens, improving generation quality and distributional coverage. Together, our results establish contextual tokens as both an interpretable view into the internal dynamics of MM-DiTs and an effective target for improving generative models.

Research source: arXiv API.

Opinion Desk

Your latest takes, front and center.

High-space editorial layout so your voice is the main event, with direct links to full blog pages.

What's the best AI model? It depends.

A practical framework for choosing AI models by workload, with live benchmark context for writing, coding, and agentic execution.

AI ModelsBenchmarksLLMsTech Strategy

OpenClaw + OpenAI: massive strategic win or expensive integration failure?

If OpenAI acquired OpenClaw, the upside could be distribution and product speed. The downside could be product overlap, antitrust pressure, and execution drag.

M&AAI PlatformsOpenAI

The AI bubble might cool hard before the real winners emerge

The hype cycle is peaking in some segments, but infrastructure and enterprise adoption suggest a rotation, not total collapse.

AI MarketVentureHype Cycle

Gold, China, and currency influence: what matters and what is overstated

China's gold strategy matters for reserves, signaling, and pricing influence, but a full gold-standard return is still unlikely in the near term.

MacroGoldChina

Why RAM prices are still high: AI data centers, supply discipline, and pushback

Memory demand from AI infrastructure is colliding with concentrated supply and disciplined production, keeping prices elevated.

HardwareMemoryData Centers

Why tech costs more now, even when manufacturing keeps improving

Better production techniques reduce unit costs, but premium positioning, bundled software value, and market anchoring keep end-user prices high.

PricingConsumer TechApple

GTA 6 delay risk: when massive hype can turn into a launch liability

Long sequel gaps can increase expectations faster than any studio can satisfy. GTA 6 could still dominate, but over-hype raises failure risk.

GamingGTA 6Hype Cycles

Live LLM Leaderboard Pulse

Snapshot of current top-performing models by intelligence and coding benchmarks.

RankModelCreatorIntelligence IndexCoding Index
1Claude Opus 5.5 (Max, Default Fallback)Anthropic57.6N/A
2Claude Sonnet 5.5 (Max, Default Fallback)Anthropic56.0N/A
3Claude Opus 5.5 (Xhigh, Default Fallback)Anthropic56.0N/A
4Claude Opus 5.5 (High, Default Fallback)Anthropic53.6N/A
5Claude Fable 5.1 (Max, Default Fallback)Anthropic53.481.6
RankModelCreatorCoding IndexIntelligence Index
1Claude Fable 5.1 (Max, Default Fallback)Anthropic81.653.4
2Claude Fable 5.1 (Xhigh, Default Fallback)Anthropic80.753.2
3Claude Fable 5.1 (High, Default Fallback)Anthropic79.151.2
4GPT-5.6 Sol (Xhigh)OpenAI78.344.0
5Claude Opus 5 (Max)Anthropic78.050.8

Latest From The Web

Live multi-source stream from Hacker News, Reddit, and DEV Community.

Auto-refresh cadence: every 15-20 minutes via server-side fetch.

Before the Alarm Screams at 3 AM: Predicting Liam's Nocturnal Hypoglycemia with Prior Labs TabPFN

I Gave 15 AI Models Proof Their Hacking Target Was a Real Company. 73% of the Ones That Noticed Told No One.

Hacktoberfest Is Coming to Nadiad, Gujarat 🚀 Official MLH Meetup at DDU, 15 Oct

Building an Offline Arduino UNO Q Cyberdeck That Identifies Birdsong and Draws Vintage Field Notes

Stop "Vibe Checking" Your AI Agents: How to Build Production Evals in 60 Minutes

Launch HN: Tensil (YC S19) – Open-Source ML Accelerators