🔬 Research Pulse
Daily Digest
May 04, 2026
🤖 AI
🧠 LLMs
1. Can Coding Agents Reproduce Findings in Computational Materials Science?
Authors: Ziyang Huang, Yi Cao, Ali K. Shargh... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: Can LLM-based coding agents reproduce end-to-end computational materials science workflows from published papers, where success requires not just coding skill but domain-specific procedural knowledge and scientific interpretation?
Summary: AutoMat is a new benchmark testing whether LLM coding agents can reproduce claims from computational materials science papers by reconstructing underspecified procedures, operating specialized scientific toolchains, and validating findings. Across multiple agent settings and foundation models, the best system achieved only 54.1% success, with failures driven by incomplete procedure recovery, methodological deviations, and execution fragility — exposing that SWE-benchmark performance does not transfer to scientific reproducibility.
Key Results: Introduces AutoMat, a benchmark of curated claims from real materials science papers requiring agents to recover underspecified procedures, navigate specialized toolchains, and validate claims. Evaluating multiple coding agent configurations across several foundation models, the best-performing setting achieved only 54.1% success rate. Error analysis revealed failures concentrate in three modes: incomplete procedures, methodological deviations, and execution fragility — with worst performance when reconstructing workflows from paper text alone.
Key Findings:
- Best-performing coding agent setting reaches only 54.1% success on AutoMat, far below SWE-bench-style performance
- Agents perform worst when workflows must be reconstructed from paper text alone (vs. when supplementary code/configs are provided)
- Three dominant failure modes: incomplete procedures, methodological deviations from established practice, and execution fragility in scientific toolchains
Technical Novelty: First reproducibility benchmark targeting computational scientific workflows (vs. generic SWE benchmarks like SWE-bench). Novelty lies in expert-curated claim-grounded tasks that jointly test procedure recovery, toolchain navigation, and scientific claim validation — rather than isolated code generation.
What's New: Shifts agent evaluation from generic software engineering to scientific reproducibility, requiring domain knowledge and claim-level reasoning rather than just passing unit tests. Expert-curated grounding in real published claims distinguishes it from synthetic or repo-mining benchmarks.
Extension Opportunities:
- Build retrieval-augmented agents that pull from materials science documentation (e.g., VASP, LAMMPS, Quantum ESPRESSO manuals) and prior reproductions to address the 'incomplete procedures' failure mode
- Develop a verification harness that catches methodological deviations early by comparing intermediate computational outputs (lattice parameters, energies) against expected ranges before full workflow execution
- Extend the benchmark to adjacent computational sciences (chemistry, biology, climate modeling) to test whether failure modes generalize, or fine-tune domain-specialized agents on materials science workflows
Replicability: Abstract does not specify code/data release; benchmark presumably to be released given the paper's purpose. Reproducing evaluations would require access to materials science simulation toolchains (DFT codes, MD packages) plus substantial HPC compute, plus API budgets across multiple foundation models.
Research Gaps:
- No quantification of how much performance gain comes from providing partial supplementary materials vs. paper text alone
- Unclear whether failures are addressable via better scaffolding/tools or require fundamentally stronger domain reasoning in base models
2. Position: agentic AI orchestration should be Bayes-consistent
Authors: Theodore Papamarkou, Pierre Alquier, Matthias Bauer... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: Where in an agentic AI system should Bayesian principles be applied — at the LLM parameter level (computationally intractable) or at the orchestration/control layer that coordinates tools, experts, and resource allocation under uncertainty?
Summary: The paper argues that while making LLMs themselves Bayesian is computationally and conceptually impractical, the orchestration layer of an agentic system is a natural and tractable home for Bayesian decision theory. It proposes practical properties and design patterns for Bayesian control that maintain calibrated beliefs over latent task variables, update them from interactions, and select actions via utility-aware policies.
Key Results: This is a position paper, so it offers no empirical benchmarks, datasets, or numerical results. Its 'proof' is conceptual: it argues that Bayesian decision theory at the orchestration layer enables (1) maintaining calibrated beliefs over task-relevant latent quantities, (2) updating beliefs from agentic and human-AI interactions, and (3) choosing utility-aware actions. The paper supports this via design patterns and illustrative examples rather than measured outcomes.
Key Findings:
- Coherent decision-making in agentic systems (tool selection, expert routing, resource allocation) is fundamentally a decision-under-uncertainty problem well-matched to Bayesian decision theory.
- Bayesian principles are more tractable and impactful at the orchestration layer than at the LLM parameter level.
- Calibrated beliefs plus utility-aware policies offer concrete leverage for human-AI collaboration, not just autonomous agents.
Technical Novelty: Prior work has explored Bayesian inference inside LLMs (in-context Bayesian behavior, Bayesian fine-tuning) and bandit-style routing. The novelty here is the explicit positional reframing: place the Bayesian machinery at the orchestration/control layer rather than inside the LLM, treating the LLM as a likelihood/observation channel while the orchestrator is the coherent belief-updating decision-maker.
What's New: Reframes the locus of 'Bayesian AI' debate: instead of asking whether LLMs can or should be Bayesian, it asserts the controller orchestrating LLMs and tools should be — a layered separation of concerns that sidesteps LLM-level intractability while preserving decision-theoretic coherence.
Extension Opportunities:
- Build a concrete Bayesian router/controller that maintains a posterior over which tool or expert to invoke given task context, and benchmark against bandit and heuristic baselines on tool-use benchmarks (e.g., ToolBench, τ-bench).
- Implement utility-aware budget allocation policies (compute, retrieval depth, expert consultation count) using POMDP-style planning at the orchestration layer over LLM subagents.
- Develop calibration evaluation harnesses for orchestration-level beliefs — measuring whether the controller's posterior over latent task variables (e.g., user intent, task difficulty) matches observed outcomes across human-AI sessions.
Replicability: No code, data, or experiments are referenced — it is a position paper. Reproducing the argument requires no compute; instantiating any proposed design pattern would require standard agentic-framework infrastructure (LangGraph/DSPy-style orchestrators) plus probabilistic programming tools (e.g., NumPyro, Pyro) for the belief layer.
Research Gaps:
- Lack of empirical validation — no benchmarks demonstrate that Bayesian orchestrators outperform heuristic or RL-based controllers in deployed agentic systems.
- Open question of how to specify and elicit utility functions and priors over latent task quantities in real human-AI workflows at scale.
👁️ Vision
1. Born-Qualified: An Autonomous Framework for Deploying Advanced Energy and Electronic Materials
Authors: Steven R. Spurgeon, Milad Abolhasani, Frederick Baddour... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: How can autonomous materials discovery overcome the 'valley of death' between lab-scale promise and industrial deployment, where lab-optimized systems fail to translate due to ignored manufacturability, cost, and durability constraints?
Summary: The paper proposes 'born-qualified' autonomous materials development, a framework that embeds manufacturability, cost, and durability constraints into the discovery loop from the start rather than treating them as downstream concerns. It identifies four enabling pillars — multi-objective metrics, causal models, modular infrastructure, and manufacturing-in-the-loop — as the path to closing the lab-to-deployment 'valley of death' for energy and electronic materials.
Key Results: This is a position/perspective paper rather than an empirical study — it does not present benchmarks, datasets, or quantitative measurements. It proposes a conceptual framework ('born-qualified' development) structured around four pillars: multi-objective metrics, causal models, modular infrastructure, and embedded manufacturing in the discovery loop. No experimental validation or numerical results are presented in the abstract.
Key Findings:
- Most autonomously discovered materials fail to deploy because optimization targets lab metrics rather than industrial viability (the 'valley of death').
- Four pillars are required to fix this: multi-objective metrics, causal models, modular infrastructure, and embedded manufacturing.
- Realizing the vision requires sustained community-wide coordination, not isolated lab improvements.
Technical Novelty: Reframes autonomous materials discovery from single-objective lab optimization to constraint-embedded co-design with manufacturing. Novel contribution is the explicit articulation of the four-pillar architecture and the 'born-qualified' framing, contrasting with prior autonomous-lab work (e.g., A-Lab, Ada) that treats deployment as a downstream concern.
What's New: Prior autonomous-science work focuses on accelerating lab-scale optimization loops; this paper reframes the entire objective by demanding industrial constraints be first-class citizens in the discovery loop, and prescribes a specific architectural recipe (causal models + modular infra + manufacturing integration) to achieve it.
Extension Opportunities:
- Build a concrete multi-objective acquisition function for Bayesian optimization that jointly optimizes lab performance, projected manufacturing cost (e.g., precursor pricing APIs), and durability proxies — and benchmark it against single-objective baselines on a public materials dataset like Materials Project.
- Develop a causal discovery pipeline (e.g., using DoWhy or PC algorithm) on existing autonomous synthesis logs to identify which process parameters causally drive scale-up failures vs. lab success, producing a reusable causal graph for a specific materials class (e.g., perovskites).
- Create a modular open-source orchestration layer (similar to BlueSky or ChemOS) that exposes standardized APIs for plugging in pilot-scale manufacturing modules alongside lab-scale autonomous platforms, with reference adapters for at least two common reactor types.
Replicability: No code or datasets — this is a perspective/roadmap paper. No compute requirements; reproducibility is not applicable.
Research Gaps:
- Lack of standardized multi-objective metrics that fuse lab performance with cost, manufacturability, and lifecycle/durability signals.
- Absence of causal (not just correlative) models linking synthesis parameters to scale-up outcomes, and missing modular infrastructure that connects autonomous labs to pilot-scale manufacturing.
🦾 ROBOTICS
1. Recovering Hidden Reward in Diffusion-Based Policies
Authors: Yanbiao Ji, Qiuchang Li, Yuting Hu... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: How can we extract reward signals from diffusion-based imitation learning policies without adversarial training, while maintaining strong imitation performance and providing structural guarantees on generalization?
Summary: EnergyFlow parameterizes diffusion policies as gradients of a scalar energy function, allowing simultaneous imitation learning and reward recovery without adversarial training. The paper proves the learned score recovers the expert's soft Q-function gradient and shows the conservative-field constraint both enables valid reward extraction and tightens generalization bounds, achieving SOTA imitation and downstream RL performance.
Key Results: The paper proves that under maximum-entropy optimality, denoising score matching recovers the gradient of the expert's soft Q-function. It formally shows constraining the learned field to be conservative reduces hypothesis complexity and tightens OOD generalization bounds, characterizes reward identifiability, and bounds how score estimation errors propagate to action preferences. Empirically, EnergyFlow achieves state-of-the-art imitation on manipulation tasks and outperforms adversarial IRL and likelihood-based alternatives for downstream RL (specific benchmark names/numbers not disclosed in abstract).
Key Findings:
- Score functions from denoising score matching recover the gradient of the expert's soft Q-function under maximum-entropy optimality
- Conservative-field constraints reduce hypothesis complexity and tighten OOD generalization bounds
- EnergyFlow achieves SOTA imitation performance on manipulation tasks and outperforms adversarial IRL and likelihood-based methods on downstream RL
Technical Novelty: Reframes the diffusion denoising field as the gradient of a scalar energy function, unifying generative action modeling and inverse RL. Unlike adversarial IRL (GAIL/AIRL) or likelihood-based IRL, it extracts rewards directly from score matching by enforcing a conservative-field constraint, yielding closed-form theoretical links to soft Q-functions.
What's New: First framework to unify diffusion-based generative action modeling with inverse RL via an energy-function parameterization, replacing adversarial IRL training with score matching while providing formal generalization and identifiability guarantees.
Extension Opportunities:
- Extend EnergyFlow to multi-task or language-conditioned manipulation by conditioning the energy function on task embeddings, then transferring recovered rewards across tasks
- Apply the conservative-field constraint to existing diffusion policy libraries (e.g., Diffusion Policy, 3D Diffuser Actor) as a drop-in regularizer to test OOD generalization gains
- Use the recovered soft Q-function as a critic to bootstrap offline-to-online RL fine-tuning on real robot hardware, comparing sample efficiency vs. AIRL/GAIL baselines
Replicability: Code is available at https://github.com/sotaagi/EnergyFlow. Compute requirements are not specified in the abstract, but diffusion policies for manipulation typically require a single modern GPU (e.g., A100/RTX 4090) for training; reproduction should be feasible at academic-lab scale.
Research Gaps:
- Abstract does not specify which manipulation benchmarks or quantitative metrics were used
- Identifiability of rewards is characterized but practical implications for transfer/multi-task settings remain unexplored
2. Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Authors: Yi Wang, Xinchen Li, Pengwei Xie... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: How can generalist VLA robot policies continue improving after deployment by leveraging fleet-scale autonomous rollouts and human interventions, given that offline pretraining data cannot capture distribution shifts, long-tail failures, and task variations encountered in the real world?
Summary: LWD is a fleet-scale offline-to-online RL framework that continually post-trains a pretrained generalist VLA policy from autonomous rollouts and human interventions collected across 16 robots. It introduces DIVL for robust value learning on sparse fleet rewards and QAM for extracting policies from flow-based action generators, achieving 95% average success across 8 real-world manipulation tasks.
Key Results: Validated on a fleet of 16 dual-arm robots across 8 real-world manipulation tasks (including semantic grocery restocking and 3-5 minute long-horizon tasks). A single generalist policy improved continually with accumulated fleet experience, reaching an average 95% success rate, with the largest gains observed on long-horizon tasks.
Key Findings:
- A single shared generalist VLA policy can be improved continually using pooled fleet experience without per-robot specialization
- Combining DIVL (distributional value estimation) with QAM (adjoint-matching policy extraction) stabilizes RL on heterogeneous, sparse-reward real-world data
- The largest improvements appear on long-horizon (3-5 minute) tasks, where offline pretraining alone is weakest
Technical Novelty: Two coupled contributions: (1) Distributional Implicit Value Learning (DIVL) for robust value estimation under sparse, heterogeneous fleet rewards, and (2) Q-learning via Adjoint Matching (QAM) — a policy-extraction method tailored to flow-based VLA action generators. Prior offline-to-online RL methods target Gaussian/diffusion policies; QAM is the novel adaptation to flow-matching action heads used in modern VLAs.
What's New: First demonstration of closed-loop, fleet-scale offline-to-online RL applied to flow-based VLA policies in the real world, with two new algorithmic primitives (DIVL, QAM) explicitly designed for the heterogeneous-fleet, sparse-reward, flow-policy setting that prior offline RL methods do not target.
Extension Opportunities:
- Apply LWD to cross-embodiment fleets (heterogeneous robot morphologies) to test whether DIVL+QAM generalize beyond identical dual-arm hardware
- Integrate active learning to prioritize which deployment trajectories to label or request human intervention on, reducing annotation cost while preserving long-horizon gains
- Extend the framework with safety-constrained RL (e.g., CMDP formulations) so policy updates from fleet data cannot regress on safety-critical behaviors during continual deployment
Replicability: Abstract does not mention code or dataset release. Reproduction would require a fleet of ~16 dual-arm manipulators, teleoperation infrastructure for human interventions, a pretrained VLA backbone, and substantial GPU compute for continual policy updates — effectively out of reach for academic labs without industrial robotics infrastructure.
Research Gaps:
- No reported analysis of safety regressions or catastrophic forgetting as the policy is continually updated from fleet data
- Generalization across embodiments and to tasks unseen during pretraining is not characterized; results are confined to one dual-arm platform and 8 tasks
3. Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation
Authors: Yajvan Ravan, Adam Rashid, Alan Yu... Published: 2026-04-30 | Citations: 0 arXiv | PDF
Research Question: How can we generate diverse, realistic, multi-modal training data for real-world robotic manipulation policies without requiring specialized teleoperation hardware or extensive real-world data collection?
Summary: Lucid-XR is an extended-reality data engine that runs physics simulation directly on XR headsets via the vuer framework, retargets human poses to robot embodiments, and amplifies collected demonstrations through a language-steerable, physics-guided video generation pipeline. Policies trained purely on this synthetic data zero-shot transfer to cluttered, poorly-lit real environments across dexterous manipulation tasks involving soft, granular, and rigid contact dynamics.
Key Results: Demonstrated zero-shot transfer of robot visual policies to unseen, cluttered, and badly lit evaluation environments after training entirely on Lucid-XR synthetic data, across dexterous manipulation tasks involving soft materials, loosely bound particles, and rigid body contact. Specific quantitative benchmarks are not provided in the abstract.
Key Findings:
- Browser/XR-native physics simulation enables internet-scale, latency-free immersive data collection without specialized hardware
- Physics-guided, language-steerable video generation effectively amplifies a small set of XR demonstrations into diverse training data
- Visual policies trained entirely on synthetic Lucid-XR data zero-shot transfer to unseen real-world scenes including soft-body, granular, and rigid-contact tasks
Technical Novelty: Combining browser/XR-headset-native physics simulation (vuer) with human-to-robot pose retargeting and a language-steerable, physics-guided video generation pipeline — eliminating specialized teleop rigs while amplifying collected demos via generative video.
What's New: Unlike prior teleop-rig-based or pure-sim data engines, Lucid-XR fuses on-device XR physics simulation, human-to-robot retargeting, and generative video amplification into a single accessible pipeline that requires only a consumer headset.
Extension Opportunities:
- Integrate bimanual/multi-agent retargeting to capture coordinated two-handed manipulation tasks within the vuer XR environment
- Extend the physics-guided video generation pipeline with domain randomization conditioned on failure-mode taxonomies to target specific policy weaknesses
- Add tactile/force-modality synthesis alongside vision so policies trained on Lucid-XR data can transfer to contact-rich assembly tasks
Replicability: Project website (https://lucidxr.github.io) is referenced; code/dataset availability is not explicitly stated in the abstract. Reproduction would require an XR headset (e.g., Quest-class), a robot arm/hand for evaluation, and GPU compute for the video generation pipeline (likely multi-GPU for diffusion-based video models).
Research Gaps:
- No reported quantitative success-rate comparisons against teleoperation-collected datasets or competing sim-to-real pipelines in the abstract
- Limited discussion of how the video generation pipeline preserves physical consistency for contact-rich and deformable-object tasks at scale
💻 COMPUTE
1. Affinity Tailor: Dynamic Locality-Aware Scheduling at Scale
Authors: Jin Xin Ng, Ori Livneh, Richard O'Grady... Published: 2026-04-30 | Citations: 0 arXiv | PDF
Research Question: How can multicore schedulers preserve microarchitectural locality (caches, branch predictors, prefetchers, LLC domains) for co-located workloads without sacrificing the work-conservation benefits of CFS-style load balancing or the rigidity of hard CPU partitioning?
Summary: Affinity Tailor is a userspace-guided kernel scheduler that assigns each co-located workload a demand-sized, LLC-compact CPU set used as a soft affinity hint rather than a hard partition. It preserves microarchitectural locality while remaining work-conserving, delivering 12%/3% per-CPU throughput gains on chiplet/non-chiplet servers in Google production.
Key Results: Deployed at Google, Affinity Tailor achieves geometric-mean per-CPU throughput gains of 12% on chiplet-based systems and 3% on non-chiplet systems over Linux CFS, plus 3-7% per-GB throughput gains from reduced memory residency due to faster execution.
Key Findings:
- Treating compact CPU sets as hints (not hard partitions) recovers locality benefits without sacrificing utilization when workloads burst.
- Locality wins are dramatically larger on chiplet-based systems (12%) than monolithic dies (3%), confirming LLC-boundary spread is a dominant cost in modern server CPUs.
- Throughput gains compound with memory: faster execution shortens residency, producing an additional 3-7% per-GB throughput improvement.
Technical Novelty: Reframing CPU sets as soft affinity hints sized dynamically to measured demand and chosen to minimize LLC-domain spread, rather than as hard cpusets/partitions or as static pinning. The split between an online userspace demand-estimating controller and a kernel that respects hints but falls back to global balancing is the specific architectural novelty.
What's New: Prior work treats the choice as binary — CFS-style global load balancing vs. strict cpuset partitioning. Affinity Tailor introduces a third regime: dynamically sized, topology-aware soft hints driven by an online demand estimator, with the kernel free to override for utilization.
Extension Opportunities:
- Port the userspace controller + affinity-hint kernel mechanism to sched_ext (BPF) so non-Google operators can deploy it without custom kernel patches.
- Extend the demand estimator to be NUMA- and memory-bandwidth-aware, jointly optimizing CPU set selection against DRAM channel contention rather than only LLC topology.
- Apply the soft-affinity-hint paradigm to GPU/accelerator scheduling or to container orchestrators (Kubernetes CPU manager) where current policies are also strict-partition vs. fully-shared.
Replicability: No code or dataset release is mentioned in the abstract; the system is described as deployed at Google on production chiplet/non-chiplet fleets, implying reproduction would require a modified Linux kernel, a userspace controller, and chiplet-class server hardware (e.g., AMD EPYC) with representative multi-tenant workloads — non-trivial outside a hyperscaler.
Research Gaps:
- No public benchmark suite or open implementation, making independent comparison against newer schedulers (e.g., sched_ext policies) difficult.
- The paper focuses on CPU/LLC locality; interaction with memory bandwidth, NUMA migration costs, and accelerator-attached workloads is left open.
2. A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator with In-Situ Regulation for Energy-Efficient Spiking Neural Networks
Authors: Shih-Hang Kao, Yang-Chan Hung, I-Wen Wang... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: How can SRAM-based in-memory computing for spiking neural networks operate reliably in the energy-efficient subthreshold regime at large scale, given that subthreshold operation amplifies process-voltage-temperature (PVT) sensitivity and degrades neuron firing-threshold accuracy?
Summary: The paper presents a 28-nm subthreshold SRAM-based in-memory computing macro for spiking neural networks that uses in-situ current sensors, distributed voltage regulators, and a programmable memory-cell firing threshold to make large-scale (1024×1304) subthreshold CIM resilient to PVT variation. A stride-tick batching schedule reduces buffer overhead by reusing inputs across timesteps. The prototype reaches 93.64% keyword-spotting accuracy at up to 1181.42 TOPS/W and 7.24 TOPS/mm^2.
Key Results: Fabricated a 28-nm CMOS prototype with 1024 wordlines, 1304 bitlines, and 128 shared neuron cells. Demonstrated 93.64% accuracy on keyword spotting, energy efficiency up to 1181.42 TOPS/W, and compute density of 7.24 TOPS/mm^2. The in-situ current sensors, distributed voltage regulators, and memory-cell-based programmable firing threshold provide PVT robustness while exploiting SNN sparsity.
Key Findings:
- Distributed in-situ regulation makes subthreshold current-mode CIM viable at 1024-wordline scale despite amplified PVT sensitivity
- A memory-cell-programmable firing threshold meaningfully improves neuron robustness over fixed-threshold designs under variation
- Stride-tick batching exploits SNN temporal redundancy to cut input-buffer overhead while preserving accuracy (93.64% on keyword spotting) at 1181.42 TOPS/W
Technical Novelty: Combines (1) in-situ current sensors with distributed voltage regulators to stabilize subthreshold current-mode CIM at unusually large array scale, (2) a programmable memory-cell-based firing-threshold neuron robust to PVT, and (3) a stride-tick batching schedule that reuses inputs across SNN timesteps to cut buffer overhead — a system-level integration not seen together in prior subthreshold SRAM-CIM SNN macros.
What's New: Prior subthreshold CIM work has been limited to small arrays because PVT drift overwhelms signal margins; this paper integrates feedback regulation directly into the macro and reframes the firing threshold as a programmable, PVT-trackable storage element, enabling subthreshold operation at a scale and efficiency previously reserved for above-threshold designs.
Extension Opportunities:
- Scale the macro to larger SNN workloads (e.g., vision tasks like DVS-Gesture or N-MNIST) and characterize how the stride-tick batching schedule's buffer savings change with deeper temporal windows
- Co-design a training-aware quantization/threshold-calibration flow that exploits the programmable memory-cell firing threshold to compensate for measured per-die PVT offsets post-fabrication
- Port the in-situ current sensor + distributed regulator scheme to other emerging non-volatile CIM substrates (RRAM/MRAM) to assess whether the PVT mitigation strategy generalizes beyond SRAM
Replicability: No code, RTL, or measurement dataset is referenced in the abstract. Reproduction requires a 28-nm CMOS tape-out (millions of dollars and months of fabrication), making full replication infeasible outside an industrial fab; partial replication via SPICE/behavioral models of the regulator and neuron loop is plausible on commodity compute.
Research Gaps:
- Evaluation is confined to keyword spotting; behavior on larger vision or temporal-event SNN workloads is unreported
- Long-term aging and cross-die yield statistics for the in-situ regulator/sensor loops are not characterized in the abstract
3. Learning Lindblad Dynamics of a Superconducting Quantum Processor
Authors: Johann Bock Severin, Malthe A. Marciniak, Rune Thinggaard Birke... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: How can one systematically select minimal yet adequate Lindblad models of quantum processors that balance physical fidelity with experimental/computational tractability, without over-fitting unjustified interaction or noise terms?
Summary: The paper introduces LIMINAL, a data-driven framework that fits nested Lindblad models to tomographic data and uses likelihood-ratio tests to decide which Hamiltonian and dissipative terms are statistically justified. Demonstrated on a 5-qubit superconducting processor, it identifies a minimal idling model (three-local Hamiltonian, two-local dissipation), recovers driven and shaped-pulse Hamiltonians, and probes hidden-qubit extensions in coupler dynamics.
Key Results: Applied LIMINAL to a 5-qubit superconducting processor and identified an idling model with three-local Hamiltonian terms and two-local dissipation, while statistically rejecting three-local dissipation via likelihood-ratio tests. Also recovered driven single-qubit Hamiltonians, reconstructed a shaped-pulse Hamiltonian without an analytic pulse ansatz, and tested hidden-qubit extensions in coupler-mediated dynamics.
Key Findings:
- Five-qubit idling dynamics are best described by three-local Hamiltonian terms but only two-local dissipation — three-local dissipation is not statistically supported
- Shaped-pulse Hamiltonians can be reconstructed from data without assuming an analytic pulse model
- Likelihood-ratio testing on nested Lindbladians enables principled hidden-degree-of-freedom detection in coupler-mediated dynamics
Technical Novelty: Combining nested Lindblad model families with formal likelihood-ratio testing on time-resolved tomographic data to provide statistically principled inclusion/exclusion of Hamiltonian and dissipative terms — moving beyond ad-hoc model selection or fixed-ansatz fits used in prior process tomography and Hamiltonian-learning work.
What's New: Unlike prior Hamiltonian-learning or process-tomography approaches that fix a model ansatz, LIMINAL treats model structure itself as a hypothesis to test, applying nested likelihood-ratio statistics to select the minimal adequate Lindbladian — bridging open-system identification with formal model selection.
Extension Opportunities:
- Scale LIMINAL beyond 5 qubits using tensor-network or neural ansatz representations of the Lindbladian to handle larger processors where full tomography is intractable
- Integrate LIMINAL into a closed-loop calibration pipeline that re-fits minimal models periodically and feeds drift parameters into pulse-shaping/error-mitigation routines
- Extend the framework to non-Markovian generators (e.g., time-convolutionless or memory-kernel models) and benchmark when likelihood-ratio tests favor non-Markovian terms over added Lindblad operators
Replicability: The abstract does not mention public code or datasets. Reproducing would require access to a multi-qubit superconducting processor with time-resolved process tomography capability plus moderate classical compute for fitting nested Lindbladians (5-qubit Liouvillian fits are tractable on a workstation).
Research Gaps:
- Scalability of nested Lindblad fitting and tomography to larger (>5 qubit) processors is not demonstrated
- Framework remains within the Markovian Lindblad regime; non-Markovian noise structures common in superconducting hardware are not addressed
⚡ ENERGY
1. Revealing the origin of XMCD in an altermagnet via three-dimensional control of spins
Authors: Daire Mallon, Zixuan Wu, Jheng-Cyuan Lin... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: Can X-ray magnetic circular dichroism (XMCD) be considered a direct signature of altermagnetism, given that XMCD requires spin-orbit coupling (SOC) and is not invariant under spin-space rotations — unlike the altermagnetic spin-splitting it is being used to validate?
Summary: The paper shows that XMCD in the g-wave altermagnet α-Fe₂O₃ originates from SOC-driven spin-direction symmetry breaking — not from altermagnetism itself — and is therefore not a clean signature of altermagnetic order. The authors model the anomalous, anisotropic XMCD with on-site Faraday tensors and exploit it to reconstruct full 3D vectorial maps of nanoscale spin textures including domain walls and topological solitons.
Key Results: Using the g-wave altermagnet α-Fe₂O₃, the authors demonstrate that XMCD is governed by spin-direction-induced symmetry breaking that altermagnetic spin groups ignore. The XMCD signal is highly anisotropic and decoupled from weak magnetic canting, and can be described by on-site Faraday tensors capturing locally uncompensated spin-orbital anisotropies. They reconstruct complete 3D vectorial maps of nanoscale spin textures in α-Fe₂O₃ thin films, including domain walls and topological solitons.
Key Findings:
- XMCD in α-Fe₂O₃ is highly anisotropic and decoupled from the weak magnetic canting, indicating it is not a direct signature of altermagnetic spin polarization
- Anomalous XMCD can be quantitatively described by on-site Faraday tensors that capture locally uncompensated spin-orbital anisotropies — a framework generalizable to other altermagnets
- The same dichroic response enables full 3D vectorial mapping of nanoscale spin textures, including domain walls and topological solitons in thin films
Technical Novelty: Introduces the on-site Faraday tensor formalism to describe anomalous XMCD in altermagnets via locally uncompensated spin-orbital anisotropies, plus a 3D vectorial reconstruction of nanoscale spin textures via spin-direction control — distinguishing SOC-driven dichroism from altermagnetic spin-splitting.
What's New: Prior work used XMCD as validation evidence for altermagnetism; this paper inverts that interpretation by showing XMCD reports on SOC-driven symmetry breaking that altermagnetic spin groups explicitly ignore, while simultaneously turning this 'limitation' into a powerful 3D vector imaging tool.
Extension Opportunities:
- Apply the on-site Faraday tensor framework to other candidate altermagnets (e.g., MnTe, RuO₂, CrSb) to reinterpret reported XMCD signatures and disentangle SOC-driven contributions from intrinsic altermagnetic effects
- Use the 3D vectorial mapping technique to characterize and engineer topological solitons/domain walls in α-Fe₂O₃ for spintronic and magnonic device prototypes (e.g., racetrack memory, magnon waveguides)
- Develop computational tools (DFT + symmetry analysis) that predict on-site Faraday tensor components for arbitrary altermagnetic materials, providing a screening pipeline for materials where XMCD is/isn't diagnostic of altermagnetism
Replicability: No code/data availability is mentioned in the abstract. Reproducing the experiments requires synchrotron access for X-ray PEEM/XMCD measurements, high-quality α-Fe₂O₃ thin film growth (likely PLD or MBE), and capability to control spin orientation in 3D — meaning specialized experimental facilities rather than commodity compute.
Research Gaps:
- Lack of a rigorous symmetry-aware framework distinguishing SOC-induced dichroic responses from genuine altermagnetic spin-splitting signatures
- Absence of techniques for full 3D vectorial imaging of nanoscale spin textures in collinear antiferromagnets/altermagnets
2. Programmable Integrated Magnonic Meshes
Authors: Piero Florio, Matteo Vitali, Valerio Levati... Published: 2026-04-30 | Citations: 0 arXiv | PDF
Research Question: How can magnonic (spin-wave) circuits move beyond isolated single elements to scalable, monolithically cascaded, programmable networks suitable for integrated RF signal processing?
Summary: The authors demonstrate the first scalable, programmable magnonic integrated circuits by using single-step direct laser writing in YIG to monolithically cascade waveguides, couplers, and phase shifters into multi-stage interferometric meshes. They validate the platform with 6-input/6-output, 7-stage networks operated without intermediate amplification, where output power and phase are programmable via external magnetic fields, opening a path toward large-scale spin-wave RF and quantum signal processing.
Key Results: Using single-step direct laser writing in yttrium iron garnet (YIG) and magneto-optical Kerr effect (MOKE) microscopy, the authors demonstrated: (1) spin-wave propagation with preserved phase coherence over hundreds of wavelengths in waveguides; (2) complete and periodic power transfer across several coupling lengths in coupled-waveguide directional couplers; (3) tunable arbitrary phase delays in phase shifters; (4) programmable splitters, frequency demultiplexers, and phase-controlled 2x2 routers with field-tunable output power/phase; and (5) integrated interferometric meshes with up to 6 inputs/outputs and 7 cascaded stages without intermediate amplification.
Key Findings:
- Direct-laser-written YIG waveguides preserve spin-wave phase coherence over hundreds of wavelengths.
- Coupled magnonic waveguides exhibit complete, periodic power transfer over multiple coupling lengths, and phase shifters achieve arbitrary tunable delays.
- Programmable cascaded devices — splitters, frequency demultiplexers, 2x2 phase-controlled routers, and 6x6 / 7-stage interferometric meshes — operate without intermediate amplification.
Technical Novelty: Prior magnonics work was limited to isolated components or short single devices; this paper introduces a single-step direct laser writing fabrication process in YIG that monolithically cascades waveguides, directional couplers, and phase shifters into multi-stage programmable interferometric meshes — eliminating the need for intermediate amplification and demonstrating complete unitary-style routing in the spin-wave domain.
What's New: It bridges a long-standing scalability gap in magnonics by combining a single-step monolithic fabrication route with full cascaded, field-programmable network architectures, mirroring the role of integrated photonic meshes but in the spin-wave domain at GHz frequencies.
Extension Opportunities:
- Scale the mesh beyond 6x6 / 7 stages toward universal N-port unitary processors (à la photonic Clements/Reck meshes) for analog matrix-vector multiplication and machine-learning inference at microwave frequencies.
- Integrate the YIG meshes with superconducting qubits or microwave resonators to enable on-chip quantum signal routing, leveraging magnon-photon hybrid coupling for quantum transduction.
- Develop closed-loop electronic control of the local magnetic-field actuators to auto-calibrate phase shifters and demonstrate trainable/reconfigurable magnonic neural networks or RF beamformers.
Replicability: The abstract does not mention public code or datasets. Reproduction requires specialized hardware: a YIG thin-film substrate, a direct laser writing system, microwave excitation antennas, electromagnets/coils for field programming, and a magneto-optical Kerr effect (MOKE) microscope — i.e., a well-equipped magnonics/spintronics lab rather than computational compute.
Research Gaps:
- No demonstration yet of fully universal N-port unitary processing or on-chip programmable training analogous to photonic deep-learning meshes.
- Insertion loss, energy-per-operation, switching speed, and integration with CMOS or superconducting electronics are not addressed in the abstract.
🔬 MATERIALS
1. Dynamical magnetotropic susceptibility as a new probe of Kitaev materials and beyond
Authors: João C. Inácio, J. Schwab, G. Rakhmanova... Published: 2026-05-01 | Citations: 0 arXiv | PDF
Research Question: How can ultra-low-frequency uniform (q=0) spin and charge fluctuations be probed experimentally and theoretically in correlated-electron systems, particularly Kitaev candidate materials like α-RuCl₃, where conventional probes struggle to access this regime?
Summary: The paper formalizes dynamical magnetotropic susceptibility k(ω) as a linear-response probe of q=0 spin and charge fluctuations and applies it to α-RuCl₃ via sign-problem-optimized AFQMC. Demonstrates that observed k(0)/T vs B/T scaling is a signature of dominant Kitaev coupling and predicts a single Larmor-frequency peak in k''(ω), positioning the cantilever technique as a sensitive discriminator between competing Hamiltonians and a probe for Kondo-destruction criticality.
Key Results: Derived dynamical magnetotropic susceptibility k(ω) via linear response theory for a generic correlated-electron Hamiltonian. Using auxiliary-field quantum Monte Carlo with ML-based sign-problem optimization, computed k(ω) for several α-RuCl₃ models and demonstrated that low-temperature k(0)/T scaling with B/T emerges only when Kitaev coupling dominates — non-Kitaev-dominant parameter sets fail to reproduce this scaling. Result is robust to optical phonon inclusion. k''(ω) shows a single peak at the Larmor frequency with local-moment features at high and low T.
Key Findings:
- k(0)/T scaling with B/T at low T uniquely signals dominant Kitaev exchange — alternative parameter sets for α-RuCl₃ fail this test
- Inclusion of optical phonons does not destroy the Kitaev-driven scaling, indicating robustness of the diagnostic
- k''(ω) exhibits a single Larmor-frequency peak with local-moment features, providing a direct dynamical fingerprint accessible to torque-cantilever experiments
Technical Novelty: First derivation of dynamical (finite-ω) magnetotropic susceptibility within linear-response theory unifying spin and charge sectors, combined with ML-optimized sign-problem mitigation in AFQMC for realistic Kitaev material Hamiltonians — prior work on k(ω) was largely static (k(0)) and phenomenological.
What's New: Unifies static and dynamical magnetotropic response in a single linear-response framework spanning insulating magnets, metals, and quantum critical points, and pairs it with ML-augmented AFQMC — bridging an experimental cantilever technique with state-of-the-art numerics for Kitaev materials.
Extension Opportunities:
- Apply k(ω) framework to Kondo destruction quantum critical points in heavy-fermion strange metals to detect uniform charge fluctuation signatures
- Extend AFQMC + ML sign-problem optimization to other frustrated magnets (e.g., triangular lattice, pyrochlores) to identify dominant-coupling fingerprints in k(0)/T scaling
- Develop experimental protocols correlating cantilever damping (k''(ω)) with predicted Larmor-frequency peak structure to discriminate between competing α-RuCl₃ Hamiltonian parameter sets
Replicability: No explicit code/data link mentioned in abstract. Reproduction requires AFQMC infrastructure (e.g., ALF, QUEST) plus ML sign-problem optimization pipeline; computationally heavy — likely O(10⁴–10⁵) CPU-hours for finite-T finite-B sweeps on Kitaev-Heisenberg-Γ Hamiltonians at sufficient lattice sizes.
Research Gaps:
- No direct experimental k''(ω) measurements yet exist for α-RuCl₃ to validate the predicted Larmor peak
- Framework's applicability to Kondo-destruction criticality is asserted but not numerically demonstrated in this work
🔥 GitHub Trending
1. bestpracticaI/kalshi-ai-trading-bot
⭐ 53 stars | TypeScript
Kalshi prediction markets trading bot algorithmic automated trading TypeScript Node.js Kalshi REST API RSA signing OpenRouter LLM CLI npm quant fintech event contracts market making exchange API bot K
algorithmic-trading automated-trading cli event-contracts exchange-api fintech
2. jmerelnyc/Photo-agents
⭐ 9 stars | Python
Autonomous self-evolving agents. Vision-grounded layered memory and self-written skills for LLM agents that operate your computer.
agent-memory ai-agents autonomous-agents computer-use llm photo-agents
3. hejun789/PhishGuard-AI
⭐ 3 stars | Python
ML-powered phishing URL detection system — Random Forest, Gradient Boosting, SVM & Logistic Regression with a cyberpunk Flask web app
cybersecurity flask machine-learning phishing-detection portfolio random-forest
4. Sunny-117/agentest
⭐ 2 stars | JavaScript
⚡️ AI 系统测试框架 — 面向 LLM/Agent 的轻量断言测试工具,AI 界的 Vitest
agent ai assertion evals evaluation llm
5. KeWang0622/nano-personal-agent
⭐ 2 stars | Python
A personal AI agent in one Python file. Persona + persistent memory + self-improving skills, all stored as plain markdown. No vector DB — the prompt cache is the index.
agent ai ai-agents anthropic claude llm
6. raghu-3113/colorectal-cancer-prediction
⭐ 1 stars | Jupyter Notebook
Machine learning project for predicting colorectal cancer risk using patient data
classification data-science healthcare-ai machine-learning python
7. AhmadUghurluzada/customer-churn-prediction
⭐ 1 stars | Jupyter Notebook
An end-to-end machine learning system that predicts whether a telecom customer will cancel their subscription.
churn-prediction fastapi gradient-boosting machine-learning python scikit-learn
8. rohanmistry231/Audit.md
⭐ 1 stars | Unknown
A structured audit framework for ML pipelines. 14 stages, 100+ checks, and every leakage pattern that silent model failures are made of.
audit-checklist data-leakage data-science feature-engineering machine-learning ml-pipeline
9. Anirodh-Padhy/documind-ai
⭐ 1 stars | Python
AI-powered multi-document RAG system with semantic search, FAISS, and conversational memory.
ai chatbot faiss llm machine-learning nlp
10. nkorvyakov28-AS/adaptersentry-m1
⭐ 1 stars | Python
Static security scanner for LoRA adapters (.safetensors) — M1 static analyzer for weight-level anomalies.
llm lora machine-learning model-security pytorch safetensors
11. Ridwan155-Akasake/cuffless-bp-estimation
⭐ 1 stars | JavaScript
Limitations of Low-Cost PPG Sensors for Cuffless Blood Pressure Estimation Using IoT and Machine Learning — undergraduate capstone (Daffodil International University)
biomedical-signal-processing blood-pressure capstone-project catboost cuffless-blood-pressure esp32
12. eyyupkln/car-price-prediction
⭐ 1 stars | Jupyter Notebook
Türkiye ikinci el araç fiyat tahmini | Turkey used car price prediction with XGBoost & FastAPI
docker fastapi machine-learning price-prediction python turkey
13. prathameshsail72-hue/sentiment-analyzer
⭐ 1 stars | HTML
An end-to-end Natural Language Processing (NLP) web application that analyzes text sentiment and provides real-time confidence scores. Built with Flask, containerized using Docker, and deployed via Go
artificial-intelligence docker flask google-cloud-run huggingface machine-learning
14. ZiyaAhmad/science-fair-2025-26
⭐ 1 stars | Jupyter Notebook
Synopsys 2025-26: PulsePath: A Transformers-GRU Hybrid Model for Monitoring SV Arrhythmias Using Wearable Arduino ECG for WPW’s Patients
arrhythmia cardiology-ai deep-learning ecg gru healthcare
15. pablo-reyes8/deepseek-v4-mini-pytorch
⭐ 1 stars | Python
From-scratch, paper-faithful PyTorch implementation of DeepSeek-V4 core architecture for transparent study, testing, ablation, and mini-scale training.
compressed-attention deep-learning deepseek deepseek-v4 from-scratch language-model
Generated by Research Pulse on 2026-05-04 06:07