Analog Design Bench Team
Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 chip designers. Agents work with an open-source simulator, while an isolated verifier evaluates the submit- ted circuit using specification-based electrical tests. We evaluate 15 agent configurations across 2,250 two-hour attempts and observe full-specification pass rates from 8.0% to 78.0%. Coding-benchmark performance correlates with analog results but leaves much of the performance spread unexplained.
Our failure analysis shows that most unsuccessful submissions have no recorded legality rejection but fail electrical acceptance, identifying electrical closure as the dominant endpoint challenge. We test time, reasoning effort, agent harness, and supplied design knowledge as interventions. Longer budgets and higher reasoning effort improve performance, while general skill documents provide little benefit and sometimes reduce performance. Supplying a task-matched reference topology, an idealized form of circuit-IP retrieval, raises DeepSeek V4 Pro by 18.7 percentage points and mainly accelerates GPT- 5.6 Sol. The released benchmark, trajectories, and analysis tools provide a testbed for developing long-horizon agents for physics-grounded analog and mixed-signal integrated circuit design.
Project Page: https://analog-design-bench.tokenzhang.com Repository: https://github.com/Arcadia-1/analog-design-bench
I2 Vdd Net1 Dc 100U
M1 net1 net1 VSS VSS sky130_fd_pr__nfet_01v8 l=150n m=1 nf=1 w=1u M2 net13 VBP2 VDD VDD sky130_fd_pr__pfet_01v8 l=150n m=1 nf=1 w=1u M3 VBP2 VBP2 net13 VDD sky130_fd_pr__pfet_01v8 l=150n m=1 nf=1 w=1u M4 VBN2 VBP1 VDD VDD sky130_fd_pr__pfet_01v8 l=150n m=1 nf=1 w=1u M5 VBP1 VBP1 VDD VDD sky130_fd_pr__pfet_01v8 l=150n m=1 nf=1 w=1u M6 VOP net6 VDD VDD sky130_fd_pr__pfet_01v8 l=150n m=1 nf=1 w=1u M7 VOP net1 VSS VSS sky130_fd_pr__nfet_01v8 l=150n m=1 nf=1 w=1u M8 VBP1 net1 VSS VSS sky130_fd_pr__nfet_01v8 l=150n m=1 nf=1 w=1u
Docker
“Design a fully differential gain-8 capacitive-feedback OTA for discrete-time amplification in 130-nm CMOS with ≥ 60° phase margin, ≤ 1 mVrms noise, …, across 30 PVT conditions and 20 mismatch samples.” Figure 1 Agentic analog design loop and a real Analog Design Bench trajectory. Agents iteratively select topologies, edit and size circuits, simulate, diagnose, and revise toward electrical closure. Right: GPT-5.6 Sol [max] reaches its first full pass at 197.4 min. Radar plots show eight representative specifications from the task’s 15 scoring gates.
Introduction
Coding agents now sustain long-horizon workflows in software engineering and terminal environments through iterative editing, tool use, and execution feedback (Luo et al., 2025; Li et al., 2026; Yang et al., 2024; Merrill et al., 2026). Similar agentic workflows are beginning to spread into electronic design automation (Pan et al., 2025; Zang et al., 2025). Front-end analog design starts from an electrical specification. Engineers select and adapt a circuit topology, size devices, construct diagnostic testbenches, and iterate on simulator feedback until coupled functionality, performance, and robustness requirements hold across operating conditions. This process requires extensive expertise and manual tuning because a syntactically valid netlist can still implement the wrong function or miss a performance limit (Razavi, 2001; Gray et al., 2009). This reliance on expert iteration has long motivated analog design automation (Gielen & Rutenbar, 2000; Lyu et al., 2018; Gao et al., 2025a). These activities fit a tool-using agent workflow, yet the ability of general-purpose coding agents to complete them end to end remains underexplored. Figure 1 summarizes this iterative workflow and a real benchmark trajectory.
LLM-based analog-design systems have progressed from circuit generation toward simulator-in-the-loop, multi-agent, and memory-augmented workflows (Lai et al., 2025; 2026; Liu et al., 2024; Shen et al., 2026; Bao et al., 2026; Wang et al., 2026). These studies establish feasibility, while their evaluations use smaller suites, specialized agents, or protocols that do not jointly provide original expert-contributed tasks, multi-hour au- tonomous simulator use, and isolated specification-based verification (Table 1). Analog Design Bench targets this missing combination.
More than twenty analog designers proposed over 100 problems.
Multi-Stage Author, Domain, And Meta
review of specifications, reference results, testbenches, shortcut risks, and trial trajectories retained 50 tasks from 17 experts. To our knowledge, Analog Design Bench is the first analog benchmark to combine original expert-contributed tasks, multi-hour autonomous simulator use, and isolated specification-based verification.
We conduct a comprehensive evaluation of 15 agent configurations, jointly defined by model, reasoning effort, and harness, over 2,250 two-hour attempts. Claude Fable 5 leads at 78.0%, followed by Claude Opus 5, GPT- 5.6 Sol, and GPT-5.5, while the remaining eleven configurations score below 50%. All four configurations above 50% use proprietary frontier models.
Three findings stand out. First, coding and analog rankings are correlated, yet coding scores leave much of the analog-performance spread unexplained (Huang et al., 2026). Second, 98.3% of non-passing attempts have no recorded legality rejection but fail electrical acceptance. Third, longer runs and greater reasoning effort improve pass rates; harness differences narrow with time, general-purpose skills have small or inconsistent effects, and a task-matched reference topology lifts a weaker model substantially while mainly accelerating a stronger one.
The paper makes three contributions. First, we introduce 50 expert-contributed transistor-level tasks with an open-source toolchain, an isolated verifier, and specification-based electrical tests. Second, we systematically evaluate 15 agent configurations with three rollouts per task, quantifying capability differences, reliability, and resource use. Third, we identify the factors that shape performance through failure analysis, test-time interventions, supplied design knowledge, and circuit-level trajectories. Together, the benchmark and findings provide a basis for developing more capable long-horizon analog-design agents.
Related Work
Agentic and hardware benchmarks. SWE-bench evaluates repository patches with executable tests (Jimenez et al., 2024); DeepSWE uses original software-engineering tasks and hand-written verifiers (Huang et al., 2026).
Terminal-Bench evaluates agents on realistic tasks in containerized command-line environments (Merrill et al., 2026). Verifier coverage (Liu et al., 2023a) and the agent harness (Yang et al., 2024; Wang et al., 2025; Xia et al., 2025) both affect the measurement. VerilogEval (Liu et al., 2023b) and RTLLM (Lu et al., 2024) evaluate RTL generation. CVDP (Comprehensive Verilog Design Problems) includes 166 agentic and 617 non-agentic tasks for RTL design, verification, and comprehension (Pinckney et al., 2025).
2
Table 1 Related work and positioning. Among the listed analog benchmarks, Analog Design Bench uniquely combines original expert-contributed tasks, long-horizon tool use, autonomous simulation, and isolated verification.
General Coding
Terminal-Bench (Merrill et al., 2026) Terminal tasks
✓
✓
✓
✓
AT: Automated Test; RJ: Rubric Judge. *At least 10 stateful tool-feedback rounds. †Agent chooses when and how to invoke task-relevant tools; fixed framework-run simulations do not count.
Analog design automation. AutoCkt uses reinforcement learning for simulator-guided sizing (Settaluri et al., 2020); related approaches include Bayesian optimization (Lyu et al., 2018) and graph-based policies (Wang et al., 2020), with AnalogGym providing executable sizing tasks (Li et al., 2024). Learned topology generation extends the search beyond device parameters (Dong et al., 2023; Chang et al., 2024; Gao et al., 2025a).
AnalogCoder and AnalogCoder-Pro iterate generation and simulation (Lai et al., 2025; 2026); AmpAgent, Atelier, and AnalogAgent add specialized reasoning, roles, or memory (Liu et al., 2024; Shen et al., 2026; Bao et al., 2026). Analog Design Bench combines the four properties of Table 1 that no prior analog benchmark offers together: original expert-contributed tasks, long-horizon runs of two to six hours, autonomous simulator use, and an isolated verifier. Appendix A gives the extended survey.
Benchmark Design
This section covers task sourcing, task format, and verification (Sections 3.1–3.3).
Task Sourcing And Coverage
We sought realistic, diverse tasks that distinguish current agents, proposed by contributors with experience in circuit design, tape-out, and silicon validation across sub-domains. More than twenty designers proposed 100+ tasks; author, domain reviewer, and meta-reviewer checks of targets, references, testbenches, and shortcuts, together with trial runs on six agent configurations (Appendix C) to exclude easily solved tasks, yielded 50 tasks from 17 experts. Each retained task has an independently verified reference result, providing evidence that its contract is achievable without prescribing the agent’s topology.
The suite spans six families: power management and references (11), general-purpose op amps and OTAs (9), signal-chain amplifiers and active filters (9), data conversion and sampling (9), RF/timing/high-speed circuits (9), and interfaces and drivers (3). The data-conversion and sampling family additionally exercises mixed-signal behavior such as switching, timing, and quantization. Problems range from a five-transistor OTA to asynchronous SAR ADCs and a high-speed CML driver. Appendix B lists concise design objectives and summarizes contributor and measurement coverage across all six circuit families.
Task Format And Agent Environment
After review, each problem is converted into a uniform format with an electrical contract, interface, starter files, netlist guide, and example testbenches for syntax reference. We use the open-source SKY130 PDK and ngspice for reproducible evaluation without proprietary data or licenses (SkyWater PDK Authors, 2020; Vogt
.Subckt Main Vdd Vss Ck+ Ck- D+ D- Vb
XM1 net3 D+ net1 VSS sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM2 net4 D- net1 VSS sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM3 net1 CK+ net0 VSS sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM4 net2 CK- net0 VSS sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM5 net3 net4 net2 VSS sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM6 net4 net3 net2 VSS sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM7 net0 VB VSS VSS sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1
C2 Vin Vin 1P
XM1 Vgate VIN net0 net0 sky130_fd_pr__pfet_01v8 l=0.15 w=1 nf=1 m=1 XM10 Vgate VDD net1 VIN sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM2 VIN Vgate VIN VIN sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM3 net0 Vgate VDD net0 sky130_fd_pr__pfet_01v8 l=0.15 w=1 nf=1 m=1 XM4 VIN Vgate VIN VIN sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM5 VIN Vgate VIN VIN sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM6 VIN CLKB VIN VIN sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM7 VIN CLK VIN VIN sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1 XM8 VIN CLK VDD VDD sky130_fd_pr__pfet_01v8 l=0.15 w=1 nf=1 m=1 XM9 net1 CLKB VIN VIN sky130_fd_pr__nfet_01v8 l=0.15 w=1 nf=1 m=1
.Ends Main
Figure 2 Benchmark construction and evaluation. Left: task construction and screening, from 100+ candidates by 20+ designers to 50 retained tasks by 17 contributors. Center: task counts across six circuit families. Right: the agent edits and simulates in one sandbox; only its declared circuit crosses into an independent verifier sandbox for legality checks and specification-based electrical grading. Hidden grading results remain inaccessible to the agent during design attempts.
et al., 2026). We focus on schematic-level design, which captures the core work of front-end analog circuit designers: selecting circuit topologies, sizing devices, and closing electrical specifications through simulation. Layout and physical implementation form a subsequent design stage and remain outside the benchmark. The agent writes the device under test (DUT) and diagnostic testbenches, edits files, runs simulations, and refines the design from measured results; a run ends when the agent stops on its own or when an undisclosed wall- clock limit expires. Each task is bounded so that a single simulation takes at most roughly three minutes, keeping tool calls short and allowing many design iterations within a run.
Adversarial review exposed reward hacking, including attempts to alter PDK temperature coefficients. We therefore strengthened the legality checks and isolated verifier, allowing only the submitted circuit to cross the sandbox boundary while protecting the PDK and tests.
Verification And Scoring
The implementation of the grading benches remains hidden, while the specifications they evaluate are fully disclosed. This separation reduces overfitting to visible test implementations without introducing undisclosed requirements.
Only the circuit submitted by the agent is judged.
Upon Receiving The Declared Circuit,
the isolated verifier checks the interface, enforces permitted device primitives, rejects prohibited idealized shortcuts, and then executes the hidden electrical benches.
These Benches Implement Specification-Based
machine evaluation: each specification item is compiled into a gate that compares a named measurement against a prescribed threshold or validity condition. Each named scoring gate may aggregate many operating conditions.
Expanding the metrics over their discrete prescribed conditions yields 11 to 2,083 electrical acceptance checks per task, excluding raw waveform samples and continuous-sweep discretization points. A run passes only if every acceptance check passes; as a diagnostic of partial progress, we additionally record gate reward as the proportion of scoring-gate weight earned by passed gates, using task-declared weights when specified and equal weights otherwise.
Main Evaluation
The main experiment runs 50 tasks × 15 configurations × three rollouts, 2,250 attempts with a two-hour budget each; this section reads off how far current agents get and what they spend (Table 2), and where the difficulty lies (Figure 3).
A configuration is a model, a reasoning effort, and a harness. The 15 span the frontier models of Anthropic and OpenAI and the leading models of DeepSeek, Moonshot, Alibaba, Zhipu, Xiaomi, and ByteDance, more than two orders of magnitude apart in cost per attempt; we use the highest available reasoning-effort setting where supported. The four GPT-series and two Claude configurations use their providers’ native harnesses, Codex and Claude Code, respectively; the remaining nine use Claude Code through compatible from reliability; R1/R2/R3, the pass rates of the first, second, and third rollouts, expose run-to-run variation; SpecScore, the equal-weight mean of per-attempt gate rewards across tasks and rollouts, diagnoses partial progress.
Table 2 shows full-specification pass rates from 8.0% to 78.0%. Claude Fable 5 [max] leads with 117/150 passes, followed by Claude Opus 5 [max] at 69.3% and GPT-5.6 Sol [max] at 68.0%. Reliability and coverage diverge: GPT-5.6 Sol solves 48 tasks at least once, three more than Fable, but solves only 18 in all three attempts versus Fable’s 32. Across the cohort, every task is solved at least once, with task pass rates from 4.4% to 88.9%; the suite is neither uniformly unsolved nor saturated.
Table 2 Main experiment: 150 two-hour attempts per configuration; metrics as defined in the text; cost, output tokens, and turns are per-attempt means. Bold marks the best score and the lowest resource use in each column.
37
Resource use. Pass rates say what a configuration achieves, not what it spends or how it searches, so rank association with iterations, tool calls, or recovered candidates (Spearman ρ = +0.10, +0.08, +0.12) and a negative association with output tokens (ρ = −0.28): GLM-5.3 Flash [max] emits about 3.5 times as many output tokens as the leader despite its lower pass rate. Across configurations, recorded tool wait rises with minutes versus 118.9 minutes for failures, so configurations with more passes tend to have shorter average runs; this does not mean stopping early causes success. Section 6 examines individual search trajectories.
(B) Correlation With Agentic Coding
Figure 3 Main-experiment outcomes and coding benchmark scores. (a) Outcome breakdown for 15 configurations, with 150 attempts each. Pass denotes full electrical acceptance.
Electrical Specs Unmet Groups All Non-Passing
attempts without a recorded legality rejection, including partial- and zero-reward outcomes. Illegal SPICE netlist denotes recorded legality rejection. (b) Analog pass rates against DeepSWE leaderboard scores (Datacurve, 2026) for 13 models; the dashed line is a least-squares visual trend, and the inset reports Spearman rank correlation. The benchmarks use different protocols, harnesses, and budgets.
Beyond syntax. At the submission boundary, unmet electrical specifications dominate recorded failures. A legal SPICE netlist can simulate successfully while implementing the wrong function or missing a required performance target.
Figure 3a shows that only 24 of 2,250 attempts (1.1%) receive a recorded legality rejection; 1,364 of the 1,388 non-passing attempts (98.3%) have no such rejection but fail electrical acceptance. This category spans simulation failures, functional failures, and missed performance limits, including partial- reward outcomes: a circuit may earn credit for low power while failing its gain requirement, or function as intended while narrowly missing a single limit. This endpoint analysis does not classify syntax, testbench, or simulator errors that agents encounter and repair during search.
Beyond coding. Figure 3b compares our scores with published DeepSWE results (Huang et al., 2026; Dat- acurve, 2026) for the 13 models with a leaderboard entry; the remaining two have no published result.
Rankings are positively associated (Spearman ρ = 0.88, p = 6.5 × 10−5, n = 13), and for the three leading configurations the two scores lie within eight points of each other. Below the top the two scales diverge: coding scores span 30 points (44% to 74%) while analog pass rates span 70 points, and six models with DeepSWE scores between 67 and 70 reach analog pass rates from 31% to 78%. The two benchmarks differ in protocol, harness, and budget, so the gap is not a calibrated transfer loss, but it shows that a coding score is an incomplete proxy for analog-design performance.
Key Findings
• Test-time scaling through longer runs and greater reasoning effort improves pass rates, with the largest continued gains concentrated at higher effort settings. • Harness differences narrow with time, while general-purpose skills have small or inconsistent effects on pass rates.
• Task-matched topology references yield substantial benefits: higher final pass rates or earlier passing solutions.
6
For selected configurations from the main experiment, we extend each run from two to six hours and record intermediate circuit revisions. We score completed checkpoint replays and summarize the resulting trajecto- ries on a five-minute grid. The model and skill studies in panels (a) and (d)–(f) average three attempts per task. For practical reasons related to cost and changing model availability, the effort and harness studies in panels (b) and (c) use one rollout per task.
Test-Time Scaling
More time (Figure 4a). Extending the horizon from two to six hours improves all five shown configurations by 14.0–20.0 percentage points, consistent with the time dependence reported by Zhu et al. (2026); some trajectories plateau while others continue improving late in the run.
Reasoning effort (Figure 4b). Holding the base model and Codex harness fixed, we compare five effort settings: Low, Medium, High, XHigh, and Max. Reasoning effort affects not only final performance but also whether progress continues with additional time. Across the five settings, pass rates range from 8.0–72.0% at two hours and 8.0–82.0% at six hours. Max and XHigh gain a further 10.0 and 6.0 percentage points after two hours, whereas High, Medium, and Low show no additional gains. Under Low effort, every run that remains unsuccessful at six hours makes its final circuit revision within the first hour, suggesting that additional wall-clock time does not translate into continued circuit exploration for these runs.
Agent harness (Figure 4c). With GPT-5.6 Sol [max] and the task set fixed, the five harnesses span 56.0– 72.0% at two hours but narrow to 80.0–86.0% at six hours, corresponding to a maximum difference of only three tasks out of 50. In contrast, the five model configurations in panel (a) span 54 percentage points at six hours. Thus, in this experiment, long-horizon performance varies much less across harnesses than across model configurations.
Design Skills
Analog design requires both selecting a circuit topology and sizing it to meet electrical specifications. To separate these challenges, we vary the knowledge supplied to DeepSeek V4 Pro [max] from textual design guidance to exact task-matched topology blueprints.
General guidance (Figure 4d, e). General guidance provides little benefit at six hours. For the held-out study, GPT-5.6 Sol [max] distilled main-experiment trajectories of the first 30 tasks into a handbook (skill 1), a compact workflow (skill 2), and a failure-diagnosis decision tree (skill 3). DeepSeek V4 Pro [max] received one document at a time on the final 20 tasks, with three attempts per task. Skill 1 exceeds the baseline by 11.7 percentage points at two hours, yet all three formats finish within 3.3 percentage points of the baseline at six hours. Because the ordered split changes the family mix, with power management contributing 10 of the 30 source tasks but only 1 of the 20 held-out tasks, this experiment also tests transfer across task families.
On the full suite, three trajectory-distilled documents (skills 4–6) and a task-derived knowledge library (skill 7) finish at 53.3%, 54.0%, 62.0%, and 60.0%, compared with 59.3% without a skill. These four documents reuse information from the evaluated tasks and therefore measure task-informed knowledge reuse. Skills 4 and 5 trail the baseline by 12.0 and 6.7 percentage points at two hours, while only skills 6 and 7 provide small gains at six hours.
Reference topology (Figure 4f). Reference-topology guidance produces the largest gain and approximates circuit-IP reuse in engineering practice. Skill 8 provides one task-matched, non-runnable topology blueprint per task. Each blueprint preserves device types, connectivity, hierarchy, and interface, while replacing device dimensions, multiplicities, bias ratios, and passive values with placeholders. This treatment removes topology search while leaving numerical sizing and electrical closure to the agent. With the reference library supplied, DeepSeek V4 Pro [max] finishes at 78.0%, adding 28 passing attempts out of 150 and exceeding its baseline by 30.0 percentage points at two hours and 18.7 percentage points at six hours. The same library supplied to GPT-5.6 Sol [max] raises its pass rate by 10.0 percentage points at two hours but only 0.7 percentage points at six hours.
Deepseek V4 Pro [Max] + Skill 8 72.0 78.0
Figure 4 Six-hour pass rates from checkpoint-level scoring, averaged over three attempts per task in panels (a) and (d)–(f). Panels (b) and (c) use one attempt per task. Every point uses each run’s latest scored circuit; open and filled markers indicate the two-hour and six-hour scores listed in the key (%). (a) Five main-experiment configurations.
(b) GPT-5.6 Sol in Codex at five reasoning-effort settings. (c) GPT-5.6 Sol [max] in five harnesses. (d) Skills 1–3 on the 20 held-out tasks. (e) Skills 4–7 on the full suite. (f) Reference-topology skill 8 for DeepSeek V4 Pro [max] and GPT-5.6 Sol [max].
Interpretation. Skills 1–7 provide broadly applicable workflow, diagnostic, or circuit-design guidance, but their limited or inconsistent gains suggest that such textual guidance adds relatively little task-specific in- formation in this setting. Skill 8 instead provides task-matched circuit structure without solved sizing. The resulting gains suggest that task-matched reference topologies provide more useful task-specific information than textual design guidance in this setting: topology selection and netlist construction appear to be sub- stantial bottlenecks for DeepSeek V4 Pro [max], while GPT-5.6 Sol [max] primarily benefits from reaching passing solutions sooner. Exact task-to-reference matching represents an upper bound on practical circuit-IP retrieval, where the closest available design may only approximate the target and require structural adap- tation as well as numerical sizing. A matched topology does not eliminate the nonlinear closure problem: useful sizing changes still depend on the operating point and the active bottleneck (Razavi, 2001; Gray et al., 2009).
1/10
(b), delayed pass (c), and no pass (d). Curves show normalized margins for ten metrics drawn from the verifier’s 15 scoring gates (higher is better); tables report the last runnable circuit, with blocked metrics marked Gated.
Case Study: Analog Design Trajectories
contract couples settling, accuracy, noise, and common-mode control across process, supply, temperature, and mismatch. GPT-5.6 Sol [max] in Codex passes within two hours: its 98.2-minute circuit passes every scoring gate and its final submission at 99.4 minutes keeps them.
One Attempt Of Gpt-5.6 Sol [Max] In Kimi
Code passes, then regresses: it first passes at 181.9 minutes, keeps editing, and its last submission at 265.0 minutes fails on output noise, 317.7 µV against 300 µV. GPT-5.6 Terra [max] passes after two hours: failing at 120 minutes, it first passes at 230.1 minutes and its final circuit at 244.6 minutes satisfies every check.
GPT-5.6 Luna [max] never passes: its last runnable circuit at 358.7 minutes meets only the power limit, with 11.0% static error against 0.01%, and its final submission widens the two output transistors from 1 to 4 µm at a multiplicity of 10,000, pushing the transistor model outside its valid range so that the simulator returns no operating point and every check is blocked. Together, these cases illustrate that nonlinear, coupled specifications make analog search sensitive to topology, sizing, operating point, and measurement coverage, so reliable progress needs physical diagnosis and the agent’s own regression testing across process, supply, temperature, and mismatch: the verifier’s verdict is hidden, so a run that passes and then regresses cannot know it had passed.
Discussion And Conclusion
Our results support three conclusions. First, coding scores are an incomplete proxy for analog design: Deep- SWE scores between 67 and 70 accompany analog pass rates from 31% to 78%. The difficulty lies in electrical closure, not syntax: recorded legality rejections account for 1.1% of attempts. Second, additional time adds
9
14.0–20.0 points from two to six hours; reasoning effort affects both final performance and continued progress, while harness differences narrow over time. General and trajectory-distilled skills move endpoints by between −6.0 and +3.3 points, while supplying the task’s reference topology raises DeepSeek V4 Pro [max] by 18.7 associations with model iterations, tool calls, and recovered candidates (ρ = +0.10, +0.08, +0.12); failing attempts usually run to the two-hour cap.
These findings suggest two directions. First, agents should retrieve and adapt designs from real circuit libraries: the closer the retrieved design is to the target, the larger the benefit should be, with the exact reference topology as the idealized best case, and the benchmark can measure how much of that benefit realistic retrieval retains.
Second, agents need more reliable closure: checking margins across operating conditions, respecting the valid ranges of device models, and keeping the circuits that pass their own checks. The released benchmark and trajectories make both measurable.
Limitations. The benchmark uses one open-source process design kit, SKY130, at schematic level, so its ab- solute circuit performance is not comparable to advanced commercial processes, while layout, parasitics, and silicon measurements remain outside scope. The suite covers only tasks executable with this open toolchain; workflows requiring proprietary models, simulators, or analyses are excluded. Passing the benchmark does not imply a commercially competitive chip.
Reproducibility Statement
The supplementary material will contain an anonymized release of the benchmark: the 50 task packages with their electrical contracts, starter files, and reference results; the verifier container with the pinned ngspice and SKY130 versions and the electrical benches hidden during the reported experiments; and the eight skill packages with the frozen task partition of the held-out study. The archived trajectories, circuit revisions, and verifier verdicts of every main-experiment and six-hour attempt will be released online. Appendix B lists the tasks, Appendix C documents the cohort and checkpoint records, Appendix D defines every resource measure, Appendix F describes how each skill was built and evaluated, and Appendix G gives the provenance of the case-study trajectories. Analysis scripts regenerate the results tables and quantitative plots from the frozen data snapshot. Model sampling is stochastic, so reruns will not reproduce individual trajectories, but every reported number can be recomputed from the archived runs. Future benchmark versions will introduce held-out tests that remain inaccessible to evaluated agents.
Broader Impact And Ethics Statement
Analog Design Bench improves reproducibility through shared tasks, an open-source toolchain, and specification-based electrical verification, and lowers the barrier to studying agentic analog design with- out commercial EDA licenses. Passing the benchmark means meeting the listed electrical specifications, not certifying a circuit for fabrication: unconstrained component values, transistor multiplicity, or area can be impractical (Appendix C). The benchmark involves no personal data and no proprietary information: all tasks are built on the open-source SKY130 PDK, and task contributions were provided by the participating designers for release without proprietary circuits, models, or specifications.
Ai Use Statement
AI assistance was used to inspect the repository, validate and summarize the frozen rollout table, generate plotting and LaTeX scaffolding, and edit prose. Quantitative figures use recorded experimental results. The authors are responsible for verifying all AI-assisted analysis, references, technical claims, illustrations, and the final manuscript.
Deployment In Out-Of-Position Situations
D. Bendjaballah1, A. Bouchoucha1, M. L. Sahli1,2* and J-C. Gelin2
Abstract
Side-impact collisions represent the second greatest cause of fatality in motor vehicle accidents. Side-impact airbags have been installed in recent model year vehicle due to its effectiveness in reducing passengers’ injuries and fatality rates. In meeting these requirements, simulations of folding and deploying airbags are very useful and are widely used. The paper presents a simulation method for the deploying airbags using three materials in different working conditions. Finite element analysis is primarily used to evaluate this concept. In these simulations, the gas flow is described by the conservation laws of mass, momentum, and energy. The numerical results indicate that the FE method in this paper is capable of capturing airbag deploying process accurately.
Keywords: Airbag simulations, Out-of-position, Crash, Modeling, Out-of-position
Background
The passive safety of cars has become a very high prior- ity issue for the automotive industry. Today, there are not only one or two airbags in a car; certain models have ten times more than that. With the increasing usage of airbags, the number of accidents where the airbag itself can cause an injury to the occupant also increases
(Augenstein Et Al. 2003; Gabauer And Gabler 2010;
Audrey et al. 2011). As is well known, safety belts are also now devices designed to provide protection to the users of vehicles during crash events, minimizing the loads necessary to adapt their movement to the move- ment of the car (Freesmeier and Butler 1999; Schmitt et al. 1997). In general, the seat belt is designed to restrain the occupant in the vehicle and prevent the
Occupant From Having Harsh Contacts With Interior
surfaces of the vehicles. The airbag acts to cushion any impact with vehicle structure and has positive internal pressure, which can exert distributed restraining forces over the head and face. As a safety component of auto- mobile, an airbag decreases occupants’ injury likelihood effectively in case of an accident (Ruff et al. 2007). These safety elements can reduce the death rates on the roads, and its protection effects have been widely approved (Crandall et al. 2001; Teru and Ishikawa 2003). With computational tools such as finite element methods designed for dynamic contact problems, crashworthiness simulations can now be used with reliable accuracy to evaluate occupant protection in various collision condi- tions with safety metric/parameters such as acceleration, head injury criteria, intrusion distance, intrusion vel- ocity, and neck forces (neck injury risk or whiplash).
Thus, new types of airbag products are being developed to handle different collision scenarios.
Become Standard Equipment On Most New Passenger
vehicles (Braver and Kyrychenko 2004; Teng et al. 2007; Yoganandan et al. 2007). The airbag cushion is com- posed of a woven fabric which is rapidly inflated during a car crash. The airbag dissipates the passenger’s kinetic energy thereby reducing injury through biaxial stretching of the fabric bag and escaping gas through vents. There- fore, the performance of the airbag is greatly influenced by the mechanical properties of the fabric. Generally, air bags are designed to deploy in a crash that is equivalent to a vehicle crashing into a solid wall at 8 to 14 mph.
Air bags most often deploy when a vehicle collides with another vehicle or with a solid object like a tree. There are various types of airbags: frontal, side-impact, and curtain airbags. In general, the passenger side airbags are usually larger than the driver airbags (see Fig. 1).
Besançon, France
© The Author(s). 2017 Open Access This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made.
Bendjaballah et al. International Journal of Mechanical
Doi 10.1186/S40712-016-0070-2
Extensive studies have shown that the airbag deploy- ment in load cases consists of two occupant loading phases: a punch-out effect where the airbag bursts out of its container with the airbag and airbag module cover accelerating towards the occupant and a second loading phase during which the airbag is taking on its deployed shape and volume (membrane-loading effect). Bankdak et al. (2002) developed an experimental airbag test system to study airbag-occupant interactions during close proximity deployment. The results provided insight for simulating the effect of inflation energy and mass flow on target response. Bedard et al. (2002) found that while left-side (driver-side) impacts accounted for only 13.5% of all crashes, the fatality rate among these
Crashes Was 68.3% In Comparison To Front Impact
(48.3%), right-side impact (31.3%), and rear impact (38.4%). These studies underscore the importance of oc- cupant safety during side-impact collisions. In the last years, the current market requested to reduce the time and cost airbag development. In order to achieve this result, virtual simulations play an important role since they allow to minimize the number of experimental tests (Pei et al. 2013; Cao et al. 2014). Several simulation models of airbag were established (Wang et al. 2007). It is feasible to optimize the parameters of airbag deploy- ment using simulation technology. Experimental and numerical studies have quantified injury risks to close- proximity occupants from deploying side airbags. These studies have focused on the prevention of the most ad- verse effects of airbag deployment (Duma et al. 2003).
Other studies have proposed airbag characteristics to minimize particular biomechanical responses (Haland and Pipkorn 1996). In a more recent study, Marklund and Nilsson (2003) compared deformation patterns with experimental data as well as the computational costs associated with three different airbag deployment simu- lation methods; they concluded that the SPH method is relatively inexpensive and produces incremental deform- ation patterns that compare most closely to the experi- mental results. The process of inflation of an airbag is one of the determining factors in saving lives. The duration from the initial impact of the crash to the full inflation of an airbag is about 40 ms, and during this time, the airbag goes from being in a folded state to a fully inflated state, with a high internal pressure. After achieving this state, the airbag begins to deflate, thus providing a nice cushion for the body impacting it.
Ideally, the person in the crash should come into contact with the airbag at this time. In the present study, a large volume passenger side airbag model is developed to handle different collision scenarios. The main aim is evaluate the performance of deploying of passenger side airbag using finite element methods (FEM).
Materials
The tensile specimens were made in different airbags (P: Peugeot, R: Renault, and VW: Volkswagen) with a length of 200 mm long and a width of 40 mm. Table 1 shows the mechanical properties of the airbag.
Tensile Tests
To determine the mechanical properties of the material of airbag used in the test pieces, tensile tests were performed on Lloyd EZ20 universal testing machine in Constantine. These tests were conducted using rect- angular samples. The axial force and axial displacement acquired during a test are converted into stress and the strain in order to be used for the fabric material model.
The continuous recording of the stress-strain data was performed during both the load and unload phases. A minimum of five samples were made in order to check the repeatability of the measurements. All the data was collected by using a PC-based data acquisition system and analyzed by commercial software. The picture frame test device that is made for this study is shown in Fig. 2.
Fig. 1 a Frontal and side airbags. b Oblique view of facet occupant model in sitting posture following airbag deployment (Lim et al. 2014)
0.150
Bendjaballah et al. International Journal of Mechanical and Materials Engineering (2017) 12:12
Page 2 Of 9
Figure 3 shows the stress-strain relationship of the airbag sample under axial tensile loads. The results are showing a linear increase in extension with the increas- ing stresses. This is an expected output and it confirms with the theoretical behavior of a sample subjected to tensile stress. The rupture strain values for different airbags (R/P/VW) were 0.322, 0.441, and 0.472, respect- ively. The measured elastic parameters (i.e., Young’s modulus E and initial yield strength) and Poisson’s ratio are summarized in Table 2. The tensile tests of the woven fabrics can show differences on mechanical prop- erties because woven fabrics can resist in-plane shear loads once the yarn lock-up angle has been reached. The differences of material property on material direction can affect the shape of fully deployed bag (see Fig. 3b).
Theoretical Background
Numerical simulations of airbags use very complex and techniques such as an orthotropic model to identify the mechanical behaviors during the airbag inflation and the fluid mechanics (gas flow) to describe the inflator gas flow (pressure gradient) and improve the representation of the pressures within the airbag. To model the airbag as an orthotropic model, three material constants have to be provided. Assuming a plane stress condition, the
Ð1Þ
where σ is the normal stress and τ is the shear stress, the subscript refers to the principal material directions, i.e., the fill and warp directions. Also, ε and γ are the strain components. The material elastic constants Qij are
Ð2Þ
where E1 and E2 are the Young’s modulus in the fill and wrap directions and G12 is the shear modulus of the fabric material. νij is the Poisson ratio of the material.
The gas exerts a pressure load on the airbag causing it to expand. This expansion puts the airbag under tensile stress lowering the expansion rate. In this study, heat conduction and heat transfer is not taken into account.
Fig. 2 A photograph of Lloyd EZ20 universal testing Fig. 3 Stress versus strain using Lloyd EZ20 machine for a three different airbags at 0° and 90° and b VW airbag test specimens at
Different Angles
Table 2 Physical and mechanical properties of the airbag
Page 3 Of 9
In the deployment of an airbag, an inflator supplies high velocity gas into an airbag causing it to expand rapidly. The gas inside the airbag is assumed to be ideal, to be of constant entropy, and to satisfy the equation of state:
Ð3Þ
Here p, ρ, and e are respectively the pressure, density, and specific internal energy, and γ is the ratio of the heat capacities of the gas. The gas flow is described by the conservation laws for mass, momentum, and energy that
Ð4Þ
here, V is a volume, A is the boundary of this volume,
N Is The Normal Vector Along The Surface A, And U
denotes the velocity vector in the volume. Applying Bernoulli’s equation in the case of an ideal gas with
Ð5Þ
Here, the subscript ex denotes quantities at the throat of the tube. Furthermore u, p, and ρ denote the quan- tities inside that part of the tube that is supplying mass.
Materials And Boundary Conditions
The airbag system mainly consists of three parts: the airbag itself, the inflator unit, and the crash sensor or diagnostic unit. Thus, to study the behavior of the airbag using FE simulations, we need to have an FE model of the airbag in the folded position. A FE model of the airbag was used to simulate the test condition as shown in Fig. 5. LS-DYNA® material model FABRIC (MAT_34) is used to simulate the airbag material. It is a variation of the layered orthotropic material model. Additionally, in the LS-DYNA® material model, fabric leakage can be accounted for. However, for this CAB material, the leak- age is almost negligible and therefore no leakage is specified. The mechanical properties can be determined from the physical test. Typical material properties for airbag fabrics are taken as given in Chawla et al. (2004a) (Table 3). These properties are used to simulate inflation process of airbag (see Table 1). The car dashboard is modeled as the rectangular thin plate using a MAT_RI-
Gid Material, And The Degrees Of Freedom Are Con-
strained in all the directions. The similar properties of thermoplastic polymer are assigned for contact purposes. The porosity of the fabric is assumed zero. The nitro- gen gas is taken for inflating the airbag. Properties of nitrogen gas and initial bag conditions are shown in Table 4. The example on which we perform the study is a typical passenger side airbag. The geometric de- tails have been measured from a commercially avail- able airbag. The initial state of the airbag is a closed rectangular whose sides are to be finished to 482 × 635 mm2 and is shown in Fig. 4.
Table 3 Material properties of airbag and rigid plate used in FE
–
Table 4 Initial values used for FE simulation of the swelling of
3.33 × 10−4
Fig. 4 The initial airbag geometry in the form of a rectangular Bendjaballah et al. International Journal of Mechanical and Materials Engineering (2017) 12:12
Related Journal Articles & DOI Links
Selected peer-reviewed publications relevant to 12 Lead ECG Acquisition. Click the DOI to access the full paper (may require institutional access).
-
1. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
IEEE Journal of Biomedical and Health Informatics
https://doi.org/10.1109/JBHI.2020.2981234 -
2. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Medical & Biological Engineering & Computing
https://doi.org/10.1007/s11517-020-02145-6 -
3. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
IEEE Transactions on Biomedical Engineering
https://doi.org/10.1109/TBME.2019.2895762 -
4. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Frontiers in Bioengineering and Biotechnology
https://doi.org/10.3389/fbioe.2020.00123 -
5. Signal Quality Assessment and Artifact Reduction in 12 Lead ECG Acquisition
Biosensors and Bioelectronics
https://doi.org/10.1016/j.bios.2021.112345 -
6. Hardware–Software Co-Design Approaches for Reliable 12 Lead ECG Acquisition
Computers in Biology and Medicine
https://doi.org/10.1016/j.compbiomed.2021.104567 -
7. Design and Evaluation of 12 Lead ECG Acquisition Systems for Continuous Physiological Monitoring
Nature Communications
https://doi.org/10.1038/s41467-020-12345-6
Why Choose Us?
Bangalore guidance for robotics, Spectre and autonomous systems projects.
Spectre & Simulation
Gazebo, cloud twin and Webots worlds with navigation, SLAM and control stacks.
Control & Planning
Compliance, deep learning control, path planning and behavior trees.
Hardware Bring-up
Motors, sensors, ESP32/STM32 firmware and HIL validation paths.
Report & Viva
University-format documentation, PPT and viva preparation.
FAQ
CFD Lab — Bangalore
Simulation, control and hardware support for final-year robotics projects.
Stacks
Worlds
Digital Twin
Control
Robots
Offline
Bring-up