OpenOpt: An Open-Source SRAM Optimizer Based
on Equivalent Circuit Model Yikai Wang , Yiheng Wu , Can Wang , Bohao Liu , Junhao Ma , Zhuohua Liu 3 ,Qinxin Mei 2
Abstract—This paper proposes a co-optimization framework to accelerate simulation , such black-box models lack
that jointly optimizes SRAM architecture and transistor sizing physical interpretability and are prone to prediction errors using equivalent circuit models. The framework simplifies in- under sparse sampling and boundary conditions, making it active SRAM cells into equivalent RC loads and static power models, achieving up to 61.4× simulation speedup while main- difficult to meet high-reliability design requirements. Prior taining high fidelity (read/write delay error <0.22%, power error analytical RC models preserve physical topology but
<1.68%). A joint search space encompassing architecture pa- derived parasitic parameters solely from individual transistor rameters and device sizing integrates seven algorithms including geometry and only captured timing behavior, neglecting power SA, PSO, Bayesian Optimization variants, and multi-objective modeling and inter-transistor coupling effects. evolutionary algorithms. Based on FreePDK45, ablation experi- ments confirm complementary gains from architecture selection The second bottleneck is the joint optimization of ar-
and transistor sizing. Among all algorithms, MOEA/D achieves chitecture and bit-cell sizing. SRAM performance depends the best Figure of Merit (8.2721), yielding 6.2% improvement not only on transistor sizing but also on architecture-level in SNM, 73.6% reduction in area, and 42.3% reduction in peak configuration (e.g., row/column counts and column-mux ra- power. The framework is publicly available at OpenOpt:URL. tio) . Traditional optimization methods typically decouple
Index Terms—SRAM, Optimization, Transistor Sizing, Equiv-
alent Circuit Models these two aspects or perform local searches in low-dimensional spaces, failing to find globally optimal solutions in the joint
I. I NTRODUCTION architecture-sizing space. Recent two-level frameworks such
as OpenACMv decompose architecture exploration and As semiconductor process nodes continue to scale down, transistor sizing into sequential stages for tractability, but such Static Random Access Memory (SRAM), which constitutes decoupling inherently limits the search to stage-wise optima the largest proportion of on-chip storage in modern SoCs and rather than the joint global optimum. Existing automation tools AI accelerators, is critical to overall chip performance, power such as OpenRAM support architecture generation but
consumption, and yield . In advanced technology nodes, lack integrated support for transistor sizing or optimization transistor scaling leads to significant increases in process algorithms. Conversely, OpenYield focuses on transistor variation, parasitic effects, and leakage current, posing severe sizing without addressing architecture-level exploration. challenges to the trade-off between PPA (Power, Performance, To address these challenges, this work proposes a large-scale
Area) and yield in SRAM design . SRAM co-optimization framework based on equivalent circuit
In traditional analog circuit design flows, SRAM design models. Built upon the high-fidelity models of OpenYield , often relies on designer experience and time-consuming man- the main contributions include: ual iterations. To address increasingly complex design spaces, simulation-based optimization methods have become main- 1) We propose an equivalent circuit model that replaces stream – . However, this approach faces two major inactive SRAM cells with compact RC loads, achieving
bottlenecks in large-scale SRAM array design. up to 61.4× speedup with <1.68% error. The first bottleneck is simulation efficiency. Accurately 2) We construct a unified platform co-optimizing archi- capturing second-order effects in advanced processes requires tecture and transistor sizing, with ablation experiments high-precision SPICE simulation. However, SRAM arrays confirming that the two stages provide complementary
contain millions of repetitive cells, and as the array size and significant gains. increases, the runtime of full-array SPICE simulation grows 3) We integrate five single-objective and two multi- super-linearly. Although surrogate models have been employed objective optimization algorithms – and release them in open-source form, providing complete toolchain This work is supported by the National Natural Science Foundation of support for SRAM design optimization.
China (NSFC) under grant No. 622041 and the Fundamental Research
Funds for the Central Universities under grant Nos. 309250106 and Through this platform, we systematically explore the joint 30924012004. ∗ Corresponding author. architecture-transistor design space of SRAM, achieving com-
prehensive optimization of stability, power consumption, area,
Optimization Algorithms
and delay. single-objective
Array level: 𝒏𝒖𝒎𝒓𝒐𝒘𝒔 , 𝒏𝒖𝒎𝒄𝒐𝒍𝒔
Simulation acceleration is critical for large-scale SRAM Transistor level: 𝑾𝒊𝒅𝒕𝒉 𝑷𝑼, 𝑷𝑫, 𝑷𝑮 , 𝑳𝒆𝒏𝒈𝒕𝒉, 𝑽𝒕𝒉
design. Neural network surrogate models achieve speedup
Evaluation Metrics
by fitting nonlinear mappings but lack physical interpretability Simulation +
𝑺𝑵𝑴, 𝑷𝒐𝒘𝒆𝒓, 𝑫𝒆𝒍𝒂𝒚, 𝑨𝒓𝒆𝒂 Acceleration
and suffer accuracy degradation under sparse sampling and 𝒎𝒊𝒏 𝑺𝑵𝑴 boundary conditions. Analytical RC equivalent models 𝑭𝑶𝑴 = 𝒍𝒈 𝒎𝒂𝒙 𝑷𝒐𝒘𝒆𝒓 ⋅ 𝒎𝒂𝒙 𝑫𝒆𝒍𝒂𝒚 ⋅ 𝑨𝒓𝒆𝒂
SPICE simulation
preserve physical topology and require no training data, but prior work derived parasitic parameters from individual Fig. 1. OpenOpt framework overview, comprising three core modules: transistor geometry, addressed only timing metrics without equivalent circuit model for simulation acceleration, optimization algorithms modeling power, and neglected leakage current of inactive (single-objective SA/PSO/CBO/SMAC/RoSE-Opt and multi-objective NSGA- cells. Since the vast majority of cells remain unselected during II/MOEA/D), and circuit generator with SPICE simulator based on OpenYield.
read/write operations and affect performance only through parasitic loading, simplifying these cells with equivalent mod- III. P ROPOSED M ETHOD els can achieve significant acceleration while maintaining The experimental workflow is illustrated in Fig. 1. OpenOpt accuracy even under extreme process corners. first designs specific test circuits, extracts key parasitic pa- SRAM optimization has been studied from both architec- rameters from SRAM cell-level simulations, and constructs
ture and transistor perspectives. Architecture-level tools such the equivalent circuit model. Subsequently, optimization algo- as OpenRAM support automated layout generation with rithms generate candidate designs within the joint architecture- configurable row/column counts and mux ratios, but do not sizing search space, and the system evaluates each candidate integrate transistor sizing or optimization algorithms. Con- via SPICE simulation. For single-objective algorithms, a com-
versely, OpenYield focuses on variation-aware transistor prehensive Figure of Merit (FoM) is calculated and fed back sizing for bitcell stability and yield without architecture-level to refine the search strategy; for multi-objective algorithms exploration. OpenACMv adopts a two-level strategy (NSGA-II, MOEA/D), the raw performance metrics (SNM, that decouples architecture search from device sizing for Power, Delay, Area) are directly provided to guide Pareto
tractability, but this sequential decomposition cannot guarantee front exploration. This iterative process continues until the joint global optimality. More broadly, simulation-based op- termination criterion is reached. timization approaches – have become mainstream for navigating complex design spaces, yet scaling them to joint A. SRAM Cell Equivalent Circuit Design architecture-sizing search with large SRAM arrays remains an We adopt a hierarchical approach: first building a compact
open challenge. 3×3 array model with a complete target cell surrounded by equivalent circuit elements, then extending it to arbitrary array
B. Problem Formulation sizes through linear superposition. As shown in Fig. 2, a
The exploration of the SRAM design space can be formu- complete 6-transistor SRAM cell is placed at the center as the lated as a constrained multi-objective black-box optimization target cell, while the surrounding 8 positions are replaced by problem. The primary objective is to identify an optimal set equivalent circuit models. The array is equipped with complete of design variables x such that PPA (Power, Performance, peripheral circuits: the row decoder drives the selected WL
Area) and stability metrics achieve Pareto optimality while to high level; the write driver transmits data through bit line satisfying process design rules and functional constraints. differential pairs; the sense amplifier detects and amplifies
Mathematically, this problem is expressed as: weak voltage differences during read operations; and the
precharge circuit precharges bit lines before each access. T min F(x) = [f (x), f (x), . . , fm (x)] Each equivalent cell model consists of four components: BL x∈Ω (1) terminal capacitance, BLB terminal capacitance, WL terminal s.t. fj (x) ≤ yj , j = 1, . , p capacitance, and static power resistance. where x represents a mixed design variable vector composed 1) Equivalent Model for BL and BLB Terminals: For cells
of discrete architecture parameters and continuous device siz- in unselected rows, WL is grounded and the access transistor ing parameters; F(x) is the objective function vector mapping is off, so the BL/BLB input current is proportional to dV /dt, key circuit metrics such as power, performance, area, and exhibiting pure capacitive behavior. As shown in Fig. 3(a), stability; and y represents the circuit design constraints. To the equivalent capacitance CBL (or CBLB ) achieves excellent
address the challenges of multi-objective trade-offs in SRAM agreement with actual simulation results. When the access circuit design, this paper integrates various types of optimiza- transistor is completely turned off, the internal storage state has tion algorithms to accommodate different search requirements. negligible impact on external behavior, allowing us to ignore
1) Design Parameters: We conduct searches on both ar- chitecture parameters and sizing parameters under capacity constraints, with the goal of determining an array organization and optimal transistor sizing—described by the number of rows r, number of columns c, multiplexing ratio µ, number of arrays na , transistor dimensions W/L, and transistor type γ—to balance dynamic performance with memory bank area. All feasible configurations satisfy the capacity constraint:
r × c × na = Capacity (5)
where Capacity represents the total number of storage bits.
This paper sets Capacity = 3 KB × 8 = 262,1 bits, and
Fig. 2. Schematic diagram of 3×3 array with equivalent circuits. The center determines na accordingly to ensure full capacity utilization. cell is a complete 6T SRAM for testing, while surrounding cells are replaced Candidate values use a sparse grid with step size of 2: r ∈ by equivalent RC models. {8, 16, . . , 512}, c ∈ {8, 16, . , 256}, and µ is selected from a set given by process and design specifications. 2) Architecture-Level Penalty Model: When the optimizer
partitions the total capacity into na sub-arrays, additional peripheral logic is required for chip-select decoding and col- umn multiplexing. We model the resulting delay and power penalties analytically, based on two physically grounded ob- (a) BL terminal (b) WL terminal Fig. 3. Equivalent-model fit for (a) BL and (b) WL terminals. Model (green) servations. matches the measured current response (blue). Chip-select delay. To address na sub-arrays, a binary se-
lection tree with ⌈log na ⌉ stages is required. Each stage internal state differences of inactive cells and significantly contributes a gate delay τcs , which is extracted once via SPICE reduce modeling complexity. simulation of a single chip-select logic gate in the standard cell 2) Equivalent Model for WL Terminal: The WL terminal library. The total chip-select overhead is therefore: exhibits more complex piecewise behavior: pure capacitive
at lower voltages, with an additional RC network above the ∆Dcs = τcs · ⌈log na ⌉ (6) access transistor threshold (Fig. 3(b)). We adopt a simplified
This logarithmic scaling reflects the standard depth of a binary
single-capacitance model CWL,simple that provides adequate decoder tree and is well-established in digital design . accuracy with superior computational efficiency.
Column multiplexer penalty. When column multiplexing is
3) Static Power Equivalent Model: After circuit stabiliza- enabled (µ > 1), a µ-to-1 pass-gate multiplexer is inserted on tion, the VDD terminal current remains essentially constant, each output column. Its propagation delay τmux and switching modeled as Rstatic = VDD /Istatic obtained from cell-level power Pmux are similarly extracted from SPICE characteriza- simulation. tion. The read path incurs both an additional delay and power 4) Array Extension: The 3×3 model extends to arbitrary due to the multiplexer:
m × n arrays through linear superposition: ∆Dmux = τmux · ⌈log µ⌉ (7)
CBL,total = (m − 1) × CBL,unit (2)
∆Pmux = (na − 1) · Pmux (8)
CWL,total = (n − 1) × CWL,unit (3)
where Eq. (8) accounts for the leakage of unselected multi- Rstatic,unit plexers across all na − 1 inactive arrays.
Rstatic,total = (4) Multi-round access penalty. Let cout denote the required
m×n−1 number of output columns (default 16). When the effective This ensures computational complexity is decoupled from output width c/µ < cout , a single access cannot fulfill the full array size, achieving significant simulation acceleration. word width, and k = ⌊cout · µ/c⌋ sequential access rounds are needed. The total delay scales linearly:
After incorporating the equivalent circuit model for simula-
tion acceleration, the focus shifts to SRAM macro design op- where Dbase includes the array read/write delay plus ∆Dcs timization through a two-level co-optimization strategy: bank- and ∆Dmux . The dynamic power accumulates proportionally level configuration under capacity constraints, and transistor- as each round activates the sense amplifiers and write drivers level sizing. independently.
Composite penalty. Combining the above terms, the cor- 2) Multi-objective algorithms: MO algorithms directly op- rected read delay is: timize the four raw metrics—SNM, Power, Delay, and Area— to approximate the Pareto front F ∗ : (i) MOEA/D decom- ′ Drd = (Drd + ∆Dcs + ∆Dmux ) × k (10) poses the problem into N scalar subproblems via weight vec- tors {λk }, each minimizing g te (x | λ, z∗ ) = maxi {λi |fi (x) −
and similarly for the write delay (without ∆Dmux ). The zi∗ |}, with neighboring subproblems sharing offspring to pro- corrected power metrics are: mote Pareto diversity. (ii) NSGA-II ranks the combined ′ Prd = Prd + ∆Pmux + Pr,dyn × k (11) parent-offspring population by fast non-dominated sorting and crowding distance, preserving well-spread solutions across the ′ Pwr = Pwr + Pw,dyn × k (12) trade-off landscape.
This penalty model effectively penalizes configurations that IV. E XPERIMENTAL R ESULTS AND A NALYSIS blindly increase na (incurring logarithmic delay overhead) or reduce array column count below cout (incurring multi-round A. Equivalent Circuit Experiments access cost), thereby guiding the optimizer toward a balanced architecture. The equivalent circuit model demonstrates significant sim- 3) Objective Function: For single-objective optimization ulation acceleration across SRAM blocks of different scales,
algorithms (SA, PSO, CBO, SMAC, RoSE-Opt), we define a as shown in Fig. 4. The speedup is particularly pronounced comprehensive Figure of Merit that balances all performance for larger blocks: medium-scale arrays achieve up to 30× metrics: speedup, and the large 256×5 configuration reaches 61.4×. This trend arises because the number of actual SRAM cells
Pmax × Dmax × Area equivalent circuit model replaces all cells outside the target
′ ′ ′ ′ row and column, maintaining essentially linear growth in the where Dmax = max(Drd , Dwr ) and Pmax = max(Prd , Pwr ). number of simulated components.
A larger FoM indicates better overall efficiency, emphasizing
low power, small area, short access delay, and large noise margin. TABLE I
S IMULATION ACCURACY COMPARISON BETWEEN FULL SRAM CIRCUITS
For multi-objective optimization algorithms (NSGA-II, (R EAL ) AND EQUIVALENT CIRCUITS (E QUIV.) OF DIFFERENT ARRAY MOEA/D), the raw performance metrics—SNM, Power, De- CONFIGURATIONS . lay, and Area—are directly provided as optimization objec- Metric 32×1 128×6 512×5
Equiv. Real Error% Equiv. Real Error% Equiv. Real Error%
tives, allowing these algorithms to explore the Pareto front Drd 5.71e-1 5.71e-1 -0.05% 1.47e-0 1.47e-0 -0.03% 4.99e-0 4.99e-0 -0.01% without scalar aggregation. Prd,avg -1.16e-0 -1.15e-0 0.50% -1.45e-0 -1.47e-0 -0.86% -4.35e-0 -4.39e-0 -0.92%
Dwr 3.58e-1 3.58e-1 0.03% 4.90e-1 4.90e-1 0.06% 1.16e-0 1.16e-0 0.22%
C. Optimization Algorithms Pwr,avg -7.29e-0 -7.38e-0 -1.27% -8.66e-0 -8.72e-0 -0.67% -2.58e-0 -2.60e-0 -0.59%
Pwr,dyn -1.98e-0 -2.00e-0 -1.28% -2.32e-0 -2.34e-0 -0.71% -6.92e-0 -6.97e-0 -0.69%
OpenOpt integrates seven algorithms in a unified loop: at Pwr,stc -2.25e-0 -2.27e-0 -0.79% -3.72e-0 -3.71e-0 0.32% -1.11e-0 -1.09e-0 1.68%
each iteration, the optimizer proposes a candidate design, evaluates it via the equivalent-circuit-accelerated SPICE sim- ulation, and updates its search strategy. 1) Single-objective algorithms: These algorithms maximize the scalar FoM (Eq. (13)): (i) SA perturbs the current design by x′ = xt + N (0, σ(Tt )) and accepts worse solutions with Metropolis probability exp(−∆f /Tt ); the temperature follows Tt+1 = αTt (α=0.98) with a restart after 5 stagnant evaluations. (ii) PSO updates each particle’s velocity via
vit+1 = wvit + c r ⊙(pi − xti ) + c r ⊙(g − xti ), combining personal best pi and global best g with Gaussian jitter and stochastic reinitialization to avoid premature convergence. Fig. 4. Simulation speedup ratios across array combinations from 32×1 to (iii) CBO fits a Gaussian Process surrogate and selects 256×512. candidates by maximizing constrained Expected Improvement EI(x)·Pr(feasible | x). (iv) SMAC uses a Random Forest As shown in Table I, the equivalent circuit model maintains
surrogate instead, natively handling the mixed continuous- high fidelity across all array scales, with read/write delay categorical design space. (v) RoSE-Opt couples GP-based errors within 0.22% and power errors within 1.68%. These Bayesian optimization with a PPO reinforcement learning results confirm that the model adequately preserves key perfor- agent that learns an adaptive sampling policy from accumu- mance metrics while providing a fast and accurate evaluation
lated evaluations. engine for iterative optimization.
TABLE II
C OMPARISON OF DESIGN PARAMETERS AND PERFORMANCE METRICS EXPLORED BY DIFFERENT OPTIMIZATION ALGORITHMS . Measure w/o Opt. MOEA/D NSGA-II SMAC Rose opt SA CBO PSO Array Size 1 × 1 6 × 3 6 × 3 6 × 3 6 × 3 6 × 3 3 × 3 3 × 3
Array Num. 10 1 1 1 1 1 2 2
Column Mux Off On On On On On On On NMOS Model VTG VTG VTG VTH VTG VTL VTL VTL PMOS Model VTG VTH VTH VTH VTH VTH VTH VTL
Pmax (mW) 3.3 1.9 1.8 1.7 1.8 2.2 1.6 1.8
min SNM (V) 0.2 0.2 0.2 0.2 0.2 0.3 0.3 0.3
FoM 7.66 8.27 8.25 8.25 8.24 8.22 8.23 8.23
(a) FoM convergence curve (b) Area-SNM comparison (c) Power-Delay (d) PPA-1/SNM Fig. 5. Overview of optimization algorithm experimental results: (a) FoM convergence curve, (b) Area-SNM comparison, (c) Power-Delay distribution, (d) Trade-off relationship between PPA and stability.
TABLE III Random Forest surrogate; (v) RoSE-Opt : Bayesian op-
A BLATION : ARCHITECTURE VS . TRANSISTOR SIZING (MOEA/D). timization with PPO agent; (vi) MOEA/D : h=17, 8 Measure S0: Orig. S1: Arch. S2: Arch.+Siz. ∆(S0→S1) ∆(S1→S2) generations, T =20; (vii) NSGA-II : population size 1,200.
Array Size 1 × 1 6 × 3 6 × 3 — —
Array Num. 10 1 1 — — Fig. 5(a) depicts the FoM convergence trajectories. While Column Mux Off On On — — NSGA-II and RoSE-Opt exhibit rapid initial convergence,
WPD (µm) 0.2 0.2 0.0 — — MOEA/D—despite a slower onset—surpasses the FoM thresh-
WPU (µm) 0.0 0.0 0.0 — — old of 8.2 after approximately 1 iterations, ultimately
L (nm) 50.0 50.0 50.4 — —
attaining the highest solution quality among all evaluated Tread (ns) 2.5 2.3 2.2 ↓9.6% ↓1.9% algorithms.
Twrite (ns) 0.7 0.6 0.6 ↓10.7% ↓4.8%
Pmax (mW) 3.3 2.0 1.9 ↓38.8% ↓5.7% The Pareto fronts are compared in Fig. 5(b)–(d). In the min SNM (V) 0.2 0.2 0.2 — ↑6.3% Area vs. 1/SNM space (Fig. 5(b)), MOEA/D consistently Area (mm ) 0.47 0.15 0.12 ↓68.5% ↓16.6% achieves the minimum area across the entire stability spec-
FoM 7.66 8.17 8.27 ↑6.6% ↑1.2%
trum, demonstrating superior search efficiency. In the Power S0: unoptimized baseline . S1: MOEA/D architecture with original sizing. S2: full vs. Delay space (Fig. 5(c)), performance differences narrow, co-optimization. though NSGA-II and RoSE-Opt show a slight advantage in the low-power region. For the PPA–Stability trade-off (Fig. 5(d)), B. Optimization Algorithm Experiments MOEA/D yields the most robust Pareto front, while NSGA-
In this subsection, we evaluate the co-optimization frame- II leverages its elitist strategy to identify high-performance work using the equivalent circuit model to re-examine the solutions in the low-PPA regime. All optimized Pareto fronts design provided in as a starting point. The optimization dominate the reference point, confirming the efficacy of the objective function is defined in Eq. (13). proposed framework.
We systematically evaluated seven optimization algorithms. Table II presents the detailed performance comparison. To ensure a fair comparison, all algorithms were allocated Under capacity constraints, excessively small arrays incur an identical budget of 1,5 SPICE simulations. The specific large area overhead due to the high number of sub-arrays, algorithmic configurations are as follows: (i) CBO : GP while excessively large arrays suffer from excessive power
surrogate with constrained EI acquisition; (ii) PSO : pop- and delay penalties. Medium-scale configurations are consis- ulation size 20, w=0.7, c =c =1.4; (iii) SA : exponential tently favored by all algorithms. Compared to the unoptimized cooling (T =1000, Tmin =10−7 , α=0.98); (iv) SMAC : baseline, the MOEA/D result achieves the best overall FoM
(8.2721): SNM improves by 6.2% (0.2 V→0.2 V), area seven integrated algorithms, MOEA/D attains the best FoM is reduced by 73.6%, read and write delays are optimized by (8.2721), yielding 73.6% area reduction, 42.3% peak power 11.3% and 14.9% respectively, and peak power is reduced by reduction, and 6.2% SNM improvement. Ablation analysis 42.3%. confirms that architecture selection contributes ∼84% of total
FoM gains, while transistor sizing provides complementary
C. Ablation Study: Architecture vs. Sizing Contributions refinements, validating the co-optimization strategy.
To disentangle the contributions of architecture selection
and transistor sizing, we conduct an ablation study whose R EFERENCES
results are summarized in Table III. Starting from the unopti- W. Gul, M. Shams, and D. Al-Khalili, “SRAM cell design challenges in
mized baseline (Stage 0), we first apply only the architecture modern deep sub-micron technologies: An overview,” Micromachines, configuration discovered by MOEA/D (Stage 1), and then vol. 13, no. 8, Art. no. 1289, 2022.
B. Narasimham et al., “Scaling Trends and the Effect of Process
apply the full co-optimization including sizing (Stage 2). Variations on the Soft Error Rate of Advanced FinFET SRAMs,” in Architecture optimization (Stage 0→1) dominates the gains. 20 IEEE Int. Reliab. Phys. Symp. (IRPS), Monterey, CA, USA, 2023, Consolidating 10 small 1 × 1 sub-arrays into 1 pp. 1–4.
D. D. Weller, M. Hefenbrock, M. Beigl and M. B. Tahoori, “Fast and
medium-sized 6 × 3 arrays reduces peripheral overhead— efficient high-sigma yield analysis and optimization using kernel density decoder depth, sense amplifiers, precharge drivers, and global estimation on a Bayesian optimized failure rate model,” IEEE Trans. routing—by nearly an order of magnitude. The cell-array Comput.-Aided Design Integr. Circuits Syst., vol. 41, no. 3, pp. 695– 708, Mar. 2022. silicon footprint drops by 68.5% (from 0.47 mm to Y. Liu, G. Dai and W. W. Xing, “Seeking the yield barrier: High-
0.15 mm ), which already accounts for 92.9% of the total dimensional SRAM evaluation through optimal manifold,” in Proc. 60th
area reduction. Read and write delays improve by 9.6% ACM/IEEE Design Autom. Conf. (DAC), San Francisco, CA, USA, 2023, pp. 1–6. and 10.7% respectively, as the shorter bit-line and word- Y. Liu and W. W. Xing, “CIS: Conditional importance sampling for yield line parasitics of the 6 × 3 tile more than compensate optimization of analog and SRAM circuits,” in Proc. 29th Asia South for the logarithmically growing chip-select overhead modeled Pacific Design Autom. (ASP-DAC), Incheon, South Korea, 2024,
pp. 386–391. in Sec. III-B2. Peak power falls by 38.8%, reflecting both B. Peng, G. Cheng, J. Qiu, R. Wang, and L. Zhang, “MOIL: An efficient the elimination of redundant peripheral activity and the more multi-objective optimization framework for SRAM cell with incremental favorable switching-capacitance profile of the enlarged tile. learning,” in 20 Int. Symp. Electron. Des. Autom. (ISEDA), 2025, pp. 632–636. SNM remains unchanged at 0.2 V, as expected for a cell- D. Challagundla, I. Bezzam, and R. Islam, “ArXrCiM: Architectural ex-
intrinsic metric unaffected by array reorganization. The net ploration of application-specific resonant SRAM compute-in-memory,” effect is a 6.6% FoM improvement attributable solely to IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 33, no. 1, pp. 179–192, 2025. architecture selection. M. R. Guthaus, J. E. Stine, S. Ataei, B. Chen, B. Wu, and M. Sarwar, Transistor sizing (Stage 1→2) provides complementary re- “OpenRAM: An open-source memory compiler,” in 20 IEEE/ACM
finement. With the architecture fixed, MOEA/D tunes the pull- Int. Conf. Comput.-Aided Des. (ICCAD), Austin, TX, 2016, pp. 1–6.
S. Shen et al., “OpenYield: An open-source SRAM yield analysis and
up/pull-down ratio (narrowing WPD from 0.2 to 0.0 µm), optimization benchmark suite,” in IEEE 43rd Int. Conf. Comput. Des. enlarges the pass-gate (WPG : 0.1 → 0.1 µm), and (ICCD), Richardson, TX, USA, 2025, pp. 167–175. switches PMOS to high-VTH, delivering an additional 16.6% Y. Zhou et al., “OpenACMv2: An accuracy-constrained co-optimization framework for approximate DCiM,” in Proc. 63rd ACM/IEEE Design area reduction, 5.7% lower peak power, 6.3% higher SNM, Autom. (DAC), Long Beach, CA, USA, 2026, pp. 1–6.
and 1.9–4.8% faster access. The larger pass-gate strengthens S. Kirkpatrick, C. D. Gelatt, and M. P. Vecchi, “Optimization by write-ability without degrading read stability owing to the simulated annealing,” Science, vol. 220, no. 4598, pp. 671–680, May 1983. compensating high-VTH PMOS pull-up. The sizing stage adds J. Kennedy and R. Eberhart, “Particle swarm optimization,” in Proc. a further 1.2% FoM improvement. IEEE Int. Conf. Neural Netw. (ICNN), Perth, WA, Australia, 1995, vol.
Complementarity of the two stages. The ablation confirms 4, pp. 1942–1948. J. R. Gardner, M. Kusner, Z. Xu, K. Q. Weinberger, and J. P. that neither stage alone can reach the final Pareto-optimal Cunningham, “Bayesian optimization with inequality constraints,” in point: architecture optimization establishes a fundamentally Proc. 31st Int. Conf. Mach. Learn. (ICML), Beijing, China, 2014, pp. more efficient floor-plan baseline (contributing ∼84% of the 937–945.
total FoM gain), while transistor sizing exploits the remaining F. Hutter, H. H. Hoos, and K. Leyton-Brown, “Sequential model-based optimization for general algorithm configuration,” in Proc. 5th Int. Conf. headroom by shaping the read/write trade-off at the cell level. Learn. Intell. Optim. (LION), Rome, Italy, 2011, pp. 507–523. Crucially, the two stages are not redundant—architecture gains W. Cao, J. Gao, T. Ma, R. Ma, M. Benosman and X. Zhang, “RoSE- come from reduced peripheral overhead, whereas sizing gains Opt: Robust and efficient analog circuit parameter optimization with
knowledge-infused reinforcement learning,” IEEE Trans. Comput.-Aided come from cell-level electrical tuning—which is why the joint Design Integr. Circuits Syst., vol. 44, no. 2, pp. 627–640, Feb. 2025. search outperforms any decoupled approach. Q. Zhang and H. Li, “MOEA/D: A multiobjective evolutionary algorithm based
FAQ
Cadence Lab — Bangalore
Simulation, control and hardware support for final-year robotics projects.
Stacks
Worlds
Digital Twin
Control
Robots
Offline
Bring-up