Abstract—Adiabatic logic reuses the energy stored on load First, output nodes are driven by a ramped power-clock wave-
capacitances through quasi-reversible switching, enabling a lower form so that the voltage drop across any conducting transistor
minimum energy consumption than conventional static CMOS. remains small throughout the charge transfer, minimising
Yet its practicality in FinFET technologies and at multi-GHz
clock rates has yet to be investigated. This work provides resistive dissipation. Second, the charge on the load is not a systematic evaluation of Positive Feedback Adiabatic Logic dissipated to ground but returned to the oscillating supply (PFAL) simulated in the TSMC 1 nm FinFET process. A set during a subsequent recovery phase. The result is an energy of PFAL standard-cell gates were realised, along with two floor that decreases with a longer ramp time, in contrast
representative combinational circuits — a 2×2 multiplier and to the frequency-independent floor of static CMOS. Among a 4-bit comparator — and compared against static CMOS logic using the energy–delay product (EDP) and the energy advantage quasi-adiabatic logic families, Positive Feedback Adiabatic metric η = ECMOS /EPFAL . Transient simulations reveal three Logic (PFAL) has demonstrated consistently robust energy sources of non-adiabatic loss: two specific to the PMOS/NMOS efficiency at frequencies up to approximately 1 GHz in planar
latch, threshold-voltage-related loss and a previously unreported CMOS , . Systematic power-clock generation studies redundant charging of the output node and one related to the have established design foundations for four-phase supply complexity of PFAL logic trees. The low-threshold Buffer/NOT cell achieves a minimum EDP of 1.23×10−2 J·s at VCLK = 0.6 V architectures . and fCLK = 7.9 GHz, while PFAL preserves an energy benefit over static CMOS of up to roughly 5× at reduced frequencies and
elevated supply voltages. A parallel-coupled quadrature voltage- Despite these advances, the viability of PFAL at deeply controlled oscillator is designed as a realistic four-phase power- scaled FinFET technology nodes and at multi-GHz operating clock generator. With this non-ideal supply, the Buffer/NOT energy stays within 2 % of the ideal sinusoidal case at 3 GHz. A frequencies remains largely unexplored. Existing studies have loading study quantifies the phase shift and amplitude reduction been limited to planar CMOS technologies with feature sizes
induced by increasing fan-out. Overall, the results provide a of 4 nm and above. The present work therefore addresses the design-oriented evaluation of PFAL in 16nm FinFET and a following research question: Do the energy-recovery benefits motivation to exploit adiabatic logic for future low-power system of PFAL persist at advanced nodes; which additional or modi- architectures. fied loss mechanisms arise in FinFET-based implementations; Index Terms—Adiabatic logic, energy recovery, energy-delay under what conditions is functional correctness preserved;
product, FinFET, low-power digital design, Positive Feedback and are these benefits sustained when a realistic, non-ideal
Adiabatic Logic, power-clock, quadrature voltage-controlled os-
cillator. power-clock supply is employed? The working hypothesis is that PFAL maintains a net energy advantage over static
E NERGY dissipation has become one of the defining
constraints in modern digital integrated circuit design.
As CMOS technology scales into the sub-2 nm regime, static
correct-operation envelope of the logic gates.
power due to subthreshold and gate-oxide leakage grows This paper presents a systematic characterisation of PFAL proportionally with device count, while dynamic dissipation in the TSMC 1 nm FinFET process. To address the question accumulates, especially in high-throughput systems. Voltage above, this work: (i) characterises a PFAL gate library and se- scaling, once the primary lever for power reduction, is increas- lected combinational circuits; (ii) maps the frequency–voltage
ingly constrained by reliability margins, noise immunity, and envelope, where they remain functionally correct and energy- the demands of high-frequency operation. The consequence advantageous over CMOS logic; (iii) identifies the dominant is a per-operation switching-energy floor, set by the 2 CVDD 2 loss mechanisms, and (iv) quantifies the energy impact of dissipation, that voltage scaling and circuit-level techniques replacing an ideal power-clock with a fully transistor-level
can mitigate but not eliminate . parallel-coupled voltage-controlled oscillator supply. The re- Adiabatic logic offers an alternative: rather than dissipating sulting characterisation maps the operating regimes, in which the energy stored on load capacitances as heat, it recycles that PFAL remains energy-advantageous over classical CMOS energy back into the power supply through quasi-reversible logic, providing a designer-oriented guide for integration of
switching. The principle relies on two mechanisms , . adiabatic logic in emerging low-power architectures.
OUT I
t t t t t
A A
(a) (b)
B B
Fig. 2. (a) General adiabatic gate and (b) example power-clock signal.
PD VSS
waveform Φ(t). Fig. 2 shows the general structure of an adia- batic gate and one period of the power-clock, which is divided Fig. 1. Circuit visualising the structure of a conventional static CMOS gate, implementing a XOR operation. NMOS transistors form a pull-down (P D) into four phases: Evaluate (E), Hold (H), Recovery (R), and network, while PMOS transistors form a pull-up (P U ) network. Based on the Wait (W). During the Recovery phase, the charge previously
input combination the output OU T is brought high (VDD ) or low (VSS ). delivered to the output capacitance is returned to the supply.
Throughout the work fCLK = 1/(t − t ) denotes the power-
II. A DIABATIC S WITCHING P RINCIPLE clock frequency. Each gate requires both a logic function F and its complement F , producing differential outputs. Multi-
A. Energy Dissipation in Adiabatic Charging
stage pipelines require four power-clocks, each leading its Conventional static CMOS logic gates, with an example predecessor by 90◦ , so that successive stages evaluate in visualised in Fig. 1 based on a XOR operation, consist of a sequence. The dual-rail structure of adiabatic logic ensures pull-up (P U ) network, utilising PMOS transistors connected that one of the complementary outputs charges and recovers between the higher potential (VDD ) and the output and a every cycle. The activity factor of an adiabatic gate, i.e., the
pull-down network, utilising NMOS transistors connected be- fraction of clock cycles during which a node switches, is thus tween the lower potential (VSS ) and the output. Outcome is equal to 1 . determined based on the input potentials (A, B and their complements) by charging or discharging a load capacitance III. PFAL G ATE L IBRARY C. Charging of this load capacitance to a supply voltage Fig. 3 exemplifies the PFAL gate structure based on a
VDD through a switch with on-resistance R dissipates a fixed XOR/XNOR operation. The core of every PFAL gate is a energy : pair of cross-coupled CMOS inverters, which form a positive ECMOS = 2 C VDD . (1) feedback latch. Two complementary logic trees, F and F , Adiabatic switching replaces the step-voltage source with a are placed in parallel with the PMOS latch transistors, each
ramped supply of ramp time T . The charging current I = tree using only NMOS devices. One tree copies the pull- CVDD /T is constant, for a linear voltage ramp, given T ≫ down network of the static-CMOS gate implementing the
RC. The energy is I 2 RT = (RC/T )CVDD 2
and can be further same function, while the other adopts the pull-up network expressed as topology by substituting PMOS with NMOS transistors. Since the trees are differential, PFAL gates additionally require RC 2 2RC complementary inputs and produce complementary outputs.
T T This dual-rail structure can be exploited to reduce gate count
From Eq. (2) the dissipated energy is a fraction 2RC of in some combinational circuits, as demonstrated in Section V T ECMOS and can in principle be lowered by increasing the with a 2×2 Multiplier. Charge recovery in PFAL proceeds ramp time , . mainly through the latch PMOS devices, since both logic trees are disabled during the Recovery (R) phase. The latch PMOS is therefore the dominant energy-recycling path.
B. Adiabatic Logic Operation The gate library characterised in this work further includes
In addition to the slow-ramp condition, quasi-reversible Buffer/NOT, AND/NAND and OR/NOR functions with cir-
operation requires two further constraints during switching, cuit schematics visible in Fig. 4. The Buffer/NOT topology identified by the authors in : (i) a transistor is not turned follows the original PFAL proposal of the authors in . on while a voltage difference exists across its terminals, and The AND/NAND and XOR/XNOR gates were constructed (ii) a transistor is not turned off while current flows through by applying the same methodology as described previously.
it. Any violation introduces a per-cycle, non-adiabatic loss By De Morgan’s laws, the AND/NAND topology realises proportional to CV 2 , which, unlike RC/T , does not decrease OR/NOR operation when its inputs are swapped with their with longer ramp time, restraining the achievable energy floor. complements: the AND output then computes A · B = A + B Losses that obey the RC/T scaling of Eq. (2) are named (NOR), while the NAND output computes A + B (OR). A
adiabatic . To realise adiabatic charging in a logic circuit, single schematic therefore implements both gate families, with the constant supply VDD is replaced by a periodic power-clock the dual-rail outputs determining the function obtained at each
Φ the Buffer/NOT circuit and its transient zoomed-in output
F F waveform in Fig. 5. The third loss is observed in circuits with A B A A more complex logic trees F and F and is described based on the XOR/XNOR circuit and its transient waveform in Fig. 6.
PMOS latch transistors, between the power-clock Φ and the
latch output nodes (Fig. 3). During the Evaluate phase, the input signals are already settled (leading the clock by 90◦ ), thus
Fig. 3. XOR/XNOR gate visualising the structure of a PFAL gate. The cross-
coupled PMOS/NMOS latch is driven by a power-clock Φ. Differential logic one logic tree is active and the other is blocked. Both PMOS trees F and F produce complementary outputs OU T and OU T , respectively. transistors begin charging their respective output nodes as Φ ramps up. The output node, whose logic tree is active, re-
Φ Φ
ceives additional current through the parallel NMOS path and
B is driven by its PMOS alone. When the faster-rising node
reaches approximately Vth,n of the cross-coupled latch NMOS,
OUT OUT AND NAND
the corresponding latch transistor turns on and discharges the slower node to ground. However, by that point, the slower node has already been partially charged, and this energy is (a) (b) not recovered but dissipated through the NMOS to ground.
The resulting redundant voltage peaks are visible in Fig. 5
Fig. 4. PFAL (a) Buffer/NOT gate and (b) AND/NAND gate. at t ≈ 2.2 [ns] and t ≈ 2.6 [ns]. Their amplitude depends on the time required for the faster node to reach Vth,n . This terminal. Table I confirms this duality by inspection: reading time is set by the supply voltage, the output capacitance, the columns under A and B as inputs, the NAND output yields and the conducting capability of the active logic tree. To the the OR function and the AND output yields NOR. Summary author’s knowledge, the asymmetric partial charging of the
of each gate’s transistor count is done in Table II. Multi-input complementary output node and its dissipation through the gate variants are obtained by extending the F and F trees in latch NMOS has not been explicitly identified as a distinct the same way as in static CMOS. loss mechanism in the PFAL literature.
IV. C OMPUTING L OSSES IN THE 1 NM P ROCESS B. Threshold Voltage Loss Cadence transient simulations of the PFAL gate library During the Recovery phase, both input signals are in the identified three non-adiabatic loss mechanisms, one of which, Wait state and neither logic tree conducts. The only available to the author’s knowledge, has not been previously reported in charge-recovery path is the PMOS latch transistor on the high-
the PFAL literature. Two of these losses produce a per-cycle side output. Its conduction condition is
CV 2 -type dissipation that does not vanish with longer ramp
time, contrary to the RC/T -scaling adiabatic losses of Eq. (2). Vsource − Vgate ≥ |Vth,p |, (3) Both originate in the PMOS/NMOS latch and are observed where Vsource = VCLK (t) and the gate is held at the com- across the entire library. Below, they are explained using plementary output, which is at ground. The PMOS therefore conducts only while VCLK (t) ≥ |Vth,p |. Charge stored at
voltages below this threshold cannot return to the supply
TABLE I
PFAL AND/NAND GATE TRUTH TABLE . and is discharged to ground through the latch NMOS in the following cycle, producing the plateau visible in Fig. 5 at A B AND NAND A B Vout ≈ |Vth,p |, at t ∈ (2.48; 2.62) [ns]. This mechanism was 0 0 0 1 1 1 0 1 0 1 1 0 previously reported in and is confirmed in this work to 1 0 0 1 0 1 persist at the TSMC 1 nm FinFET node.
1 1 1 0 0 0 The threshold voltage loss motivates the use of low- threshold-voltage (LVT) devices in the gate library: a lower
VALID OUTPUT WITHIN ONE QUARTER OF A POWER - CLOCK CYCLE AFTER
return to the supply, while the correspondingly lower Vth,n THE INPUT IS APPLIED , CORRESPONDING TO A PER - GATE reduces the time the faster node needs to trigger the latch,
COMPUTATIONAL LATENCY OF TCLK /4. thus decreasing the redundant charge on the complementary
output. However, lower threshold voltage results in increased
XOR/XNOR 2 1 1 cycle, increasing the net per-cycle energy, even though the
Normalised output, V out/V th,p intermediate node
Voltage [V]
2.2 2.3 2.4 2.5 2.6 0.1 0.2 0.3 0.4 0.5
Fig. 5. PFAL Buffer/NOT state transitioning output waveforms. Both traces
show recovering output, with the threshold-voltage plateau starting to create Fig. 6. XOR/XNOR gate node waveforms showing charge redistribution at Vout ≈ |Vth,p | and a redundant-charging peak preceding latch resolution. from the intermediate logic-tree node (purple) driven by input N OT A (blue), into the XOR output (green) during Recovery. The resulting “knee” extends the non-recoverable voltage plateau beyond the |Vth,p | level observed in the adiabatic losses are smaller . Simulation of the Buffer/NOT Buffer/NOT gate.
gate (Section VII, Fig. 11) shows that LVT devices reduce energy consumption for operating frequencies above approxi- power-clock phase: each gate introduces a one-phase latency mately 1 MHz across the characterised supply range, while (TCLK /4), so multi-stage pipelines require careful phase as- standard-threshold-voltage (SVT) devices are preferable below signment across all four clock phases. If a signal is reused in a this region. later stage, a Buffer/NOT gate is inserted as a phase-alignment
delay element, costing energy and area. All circuits were veri- C. XOR/XNOR Charge Redistribution fied for functional correctness under trapezoidal, sinusoidal, In circuits with deeper logic trees such as the XOR/XNOR and triangular power-clock waveforms. In all schematics, gate, the F and F logic trees contain intermediate nodes PFAL gate symbols adhere to the following convention: the
between series connected NMOS transistors that accumulate upper two terminals carry A, B and the lower two carry A, charge during the Evaluate (E) phase. During Recovery (R) B. The output without a circle symbol is the positive result phase, when the input-dependent recovery paths are not fully and the one with the circle is its complement. active, the stored charge on these intermediate nodes cannot be
returned to the supply and is trapped. During the Wait phase, A. 2×2 Multiplier when inputs rise again, a path between trapped charge and the output may open (depends on the input combination) and A schematic of the PFAL 2×2 multiplier is shown in Fig. 7. the charge redistributes into the output node, extending the The circuit architecture was adapted from for the PFAL threshold voltage plateau and increasing the non-recoverable topology used in this work and is organised in two power-
energy per cycle. clock phases. For 2-bit operands A = a a and B = b b , the Fig. 6 illustrates this effect: after the output (green) reaches 4-bit product M = M M M M is computed as: the |Vth,p | plateau, input N OT A (blue) rises high and charge M = a b , (4) from the intermediate node (purple) flows into the output,
M = (a b ) ⊕ (a b ), (5)
producing a visible “knee” that delays the final discharge. This additional loss mechanism, creates a trade-off between M = (a b ) ⊕ (a a b b ) ⇔ (a b ) · (a b ), (6) the complexity of a logic tree favoring speed (output re- M = (a b ) · (a b ). (7) solves in TCLK /4), the efficiency of energy recycling and maximum operating speed. Furthermore, combined with the The key simplification over a standard CMOS implementation
larger total capacitance of the 12-transistor XOR/XNOR gate, is in Eq. (6): the complemented partial product a b is explains its reduced functional operating region compared to available directly from the complementary output of the first the Buffer/NOT and AND/NAND gates observed in the EDP AND/NAND gate (Section III), eliminating one XOR gate. characterisation of Section VII. The circuit uses 6 AND/NAND gates, 1 XOR/XNOR gate,
and 1 Buffer/NOT gate for phase alignment of M , overall V. C OMBINATIONAL L OGIC C IRCUITS 8 gates, 6 transistors (1 PMOS, 5 NMOS). The output is valid after two power-clock phases (1/(2fCLK )). The gate library and loss analysis of the preceding sec- Functional correctness is verified against a reference truth tions establish the foundation for constructing more complex table across all 1 input combinations and no mismatches were
PFAL circuits. Two combinational circuits were implemented found. and verified in TSMC 1 nm: a 2×2 multiplier and a 4-bit comparator.
A design constraint distinguishes PFAL combinational cir- B. 4-bit Comparator
cuits from their static CMOS counterparts. Since the gates The PFAL 4-bit comparator, shown in Fig. 8, determines are clocked, every stage must be supplied by the correct whether A > B, A = B, or A < B for two unsigned
Φ Φ Φ Φ
a a0b b a0b M A
PFAL PFAL
a AND NOT
4 A>B
a0b M B b a0b a a1b a1b M A b PFAL a0b PFAL a B
AND a1b XOR 4
b a1b a0b M A a a0b a1b M A<B b B
PFAL a0b PFAL
a AND a1b AND A 5 b a0b a0b M B a a1b a0b M b PFAL a1b PFAL A a B
AND a0b AND
b a1b M A 4 A=B a1b B A B Fig. 7. 2×2 PFAL multiplier. Φ : partial-product AND gates; Φ : XOR, 2 A
Fig. 8. Simplified schematic of a 4-bit PFAL comparator. Gate symbols show
4-bit operands. The circuit topology was adapted from the gate function used. The complements are routed accordingly. Integer on a for the PFAL library described in Section III, with input- gate symbol indicates the number of inputs. Phase assignments are indicated polarity assignments, phase-alignment buffers, and per-phase by Φ –Φ groupings. gate placement specified by this work. In Fig. 8 gate symbols show the used gate function and the complements are routed
redistribution (Section IV). accordingly. Integer on a gate indicates the number of inputs.
Section III. The circuit is organised across four power-clock
phases as follows: A real power-clock must store and release energy on ev- • Φ : 8 Buffer/NOT gates for signal phase extension, and ery cycle while maintaining quadrature across four phases.
4 XOR/XNOR gates to detect per-bit equality (ai ⊕ bi ). An LC resonant tank naturally fulfills the energy-storage
• Φ : 4 AND/NAND gates (2-, 3-, 4-, and 5-input) to form requirement, its sinusoidal output is simple to implement and weighted magnitude-comparison terms, incorporating the sustain by introducing transconductance acting as compensa- bit-priority hierarchy, plus 1 AND/NAND gate (4-input) tion for internal tank losses. A parallel-coupled quadrature to combine previous equality terms and produce A = B voltage-controlled oscillator (P-QVCO), consisting of two
output. NMOS-coupled, complementary cross-coupled LC oscillators, • Φ : 1 OR/NOR gate (4-input) to produce A > B is therefore selected as the power-clock topology . The and 1 Buffer/NOT gate to extend preceding 4-input NMOS coupling transistors enforce quadrature between the AND/NAND gate, thus A = B output in phase. four output phases Q+, Q−, I+, I−, thus satisfying the four-
• Φ : 1 NOR/OR gate (2-input) to produce the final A < B phase clocking requirement of Section II. output, and 2 Buffer/NOT gates to extend Φ outputs for phase alignment. A. Components and Quadrature Tuning The complete gate inventory is: 1 Buffer/NOT, Fig. 9 shows the oscillator schematic. Each LC core uses
4 XOR/XNOR (2-input), 5 AND/NAND (1×2-, 1×3-, 2×4-, two off-chip inductors (L = 2 nH, Q = 20), as the TSMC
1×5-input), and 2 OR/NOR (1×2-, 1×4-input), totalling 1 nm PDK does not include an inductor component. The
2 gates and 1 transistors (4 PMOS, 1 NMOS). required tank capacitance, determined by the target oscillation
Multi-input gates are constructed by extending the F and F frequency fosc ≈ 3 GHz, was found to be Cosc ≈ 1.4 pF (Eq. logic trees as described in Section III. Since the design spans (10)). The transistors were sized using the gm /Id method all four phases, the output is valid one full clock period TCLK targeting minimum energy consumption. The complete sizing after the inputs are applied. procedure is documented in Appendix A. The coupling-width
Functional correctness was verified using a set of 1 input ratio m = Wcoupling /WLC tank governs the trade-off between pairs that test each decision path: both A = B corners (all- waveform purity and quadrature accuracy , both important zero and all-one operands), four A > B cases in which each characteristics of adiabatic power-clocks. bit position in turn acts as the deciding bit, and the four A < The NMOS coupling transistors (M –M in Fig. 9) en-
B symmetric cases obtained by swapping the operands. All force quadrature between the two LC cores. In this work, tested combinations produced correct results. The architecture m = Wcoupling /WLC tank was swept from approximately is scalable to N -bit comparators by extending the multi-input 0.4 to 1.2. At m ≈ 0.4 the simulated phase deviation AND and OR gates. Although, the recycling efficiency was from 90◦ was approximately 15◦ between the Q and I cores,
observed to decrease as the logic trees become more complex, increasing for lower values of m. Those values were not due to higher path resistance and intermediate-node charge investigated further due to simulation time constraints. Above
M M M M
IN Losc Rs,L Rs,L Losc Losc Rs,L Rs,L Losc PFAL PFAL PFAL CMOS Q+ Q- I+ I- CL
Cosc Cosc Cosc Cosc
I- M M M M I+ Q+ M M M M Q- Fig. 10. Testbench structure for PFAL gate and combinational circuit characterisation. The DUT is preceded by two PFAL buffers (Φ , Φ ) and loaded with a minimum-sized CMOS inverter and capacitor.
Fig. 9. P-QVCO power-clock schematic. Two complementary cross-coupled
LC cores are coupled through NMOS transistors M –M . where the integration window [t , t ] spans a complete set of all possible input combinations. The total energy is normalised m ≈ 1.2 the sinusoid distortion became significant and the by the number of input-combination sets, averaging out input- oscillator power increased significantly (clarified further in dependent variations arising during adiabatic operation, or by the section). The ratio was therefore treated as a design
the number of clock periods. The figures of merit are: the knob rather than fixed to a single value, allowing the energy energy-delay product EDP = E/fCLK [J·s], its minimum measurements of Section VII to capture the sensitivity of
PFAL efficiency to waveform shape and phase offset. The
the energy gain η = ECMOS /EPFAL , also called the Energy coupling ratio m ≈ 0.4 was in the end selected as it
Saving Factor . The outputs are verified against reference
yielded the lowest energy dissipation for the Buffer/NOT logic values. An output was considered incorrect if it mis- and AND/NAND gates in the measurements of Section VII, matched the reference output or if the residual distortion from while still maintaining adequate phase differences for four-
Section IV propagated into the subsequent clock cycle with an
phase operation. Additionally, it provided a significantly lower amplitude exceeding 5 % of VCLK,max , as such a level risks overall power consumption of the oscillator comparing to corrupting the following evaluation. For the 4-bit Comparator, the largest, tested m (the largest, tested m = 1.2 dissipated the possible 2 input-output combinations made exhaus- approx. 45% more power). This is due to increase in width, of tive simulation impractical. The circuit was tested against
coupling transistors, by over a factor of 2, what significantly a representative set of combination pairs (Section V). For increases the oscillator capacitance and drawn current. the combinational circuits, the CMOS reference energy was
An NMOS-plus-PMOS coupling topology was also eval-
estimated by summing the individually measured energies of uated, while it produced a visibly cleaner sinusoid and bet- the constituent static CMOS gates. This model neglects inter- ter quadrature, the additional PMOS coupling transistors in- gate loading effects or gate-specific load variations. Therefore, creased the oscillator power dissipation considerably (by 77% it approximates the true CMOS energy consumption. All at m = 0.737). Thus, the idea was not pursued. energy comparisons assume an activity factor of 1 for both
PFAL and CMOS. In practice, lower CMOS activity factors
mately 3 GHz, the output peak-to-peak voltage is 5–8 % below would reduce the effective gain. In all EDP and energy- the DC supply voltage, and the phase spacing is approximately gain figures, the white region marks combinations where the 75◦ between Q and I cores for m = 0.474. The oscilla- gate fails the correctness criterion. Its boundary defines the tor power consumption is 8 µW at VDD = 8 mV and functional operating limit. The white dots mark the exact 13 µW at VDD = 9 mV both with m = 0.474. The impact
simulated operating point. EDP is shown on a logarithmic of the real power-clock supply on PFAL gate energy and the scale and gain on a linear scale. driving limits of the oscillator are quantified in Section VII.
All gates and circuits presented in the preceding sections
and SVT implementations of the PFAL Buffer/NOT gate are implemented and simulated in Cadence Virtuoso using the across the characterised frequency and supply range. LVT TSMC 1 nm FinFET (TSMC16ADFP) process design kits, devices extend the PMOS recovery window and reduce the with transistors based on the BSIM-CMG model . Test- latch-ambiguity time (Section IV), yielding lower energy benches for each gate and circuit follow the structure shown in consumption and a wider functional frequency range above
Fig. 10: the device under test (DUT) is preceded by two PFAL
approximately 10–1 MHz, depending on the supply voltage.
Below this crossover, subthreshold leakage dominates and
phase (Φ → Φ → Φ ), to provide a realistic adiabatic input
SVT devices become preferable . All subsequent results
waveform. The DUT output is loaded with a minimum-sized use LVT, consistent with the GHz target of this work.
Fig. 1 presents the EDP heatmap of the LVT PFAL
node. The energy drawn from the power-clock is recorded as:
Z t
frequency. Within the correct operation envelope, the mini- E= VCLK (t) ICLK (t) dt, (8) mum EDP is consistently found at the lowest supply voltage t
1 VCLK =0.4 V VCLK =1.0 V 0.8 5
LVT preferred
VCLK =0.9 V 2
Normalised EDP
2 0.5 1 1 1 0.4
0.3 5 0.5 5 6 7 8 9 1 3
1 1 1 1 1 1
Frequency [Hz]
10-1 1 Fig. 11. Energy per period ratio of ELVT /ESVT indicating preferred device 0.3 0.4 0.5 0.6 0.7 0.8 0.9 type for a given operating point (left) and operating points where only LVT VCLK [V]
devices are functional (right) for different power-clock voltages VCLK .
Normalised EDP
1 3
5 2
Normalised EDP
3 1
10-1 1 0.3 0.4 0.5 0.6 0.7 0.8 0.9 VCLK [V] 2
Fig. 12. EDP per period heatmap of the LVT PFAL Buffer/NOT gate in 1 nm with an ideal trapezoidal power-clock. The red dot marks the minimum EDP 10-1 1 0.3 0.4 0.5 0.6 0.7 0.8 0.9 of 1.2 × 10−2 J·s at VCLK = 0.6 V and fCLK = 7.9 GHz. VCLK [V]
XOR/XNOR gate in 1 nm with an ideal trapezoidal power-clock. The red
dot marks the minimum EDP of 1.0 × 10−2 J·s at VCLK = 0.8 V and corresponding to the highest achievable frequency. This is fCLK = 7.9 GHz. expected, as PFAL gate energy is robust against supply-voltage variations, while the 1/fCLK factor in the EDP decays with increasing frequency. Therefore, the frequency term dominates B. 1 nm PFAL vs. CMOS with Ideal Supply
the product and drives the minimum toward the fastest feasible
Fig. 1 presents the energy gain η plot of the LVT PFAL
operating point. For lower frequencies, the heatmap shows
Buffer/NOT gate over its static CMOS inverter equivalent
that reducing the supply voltage maintains a comparable EDP at a given operating point. The PFAL circuits have energy level, providing a practical voltage-selection guide for a given advantage over CMOS across the whole available operating frequency. region. The energy gain has the highest magnitude for lower Fig. 13, 1 present the EDP heatmaps of the LVT PFAL frequencies, due to the 1/T scaling as shown in Eq. 2, and
AND/NAND and XOR/XNOR gates as a function of supply for larger supply voltages, due to the fact that CMOS gates voltage and operating frequency respectively. Both heatmaps dissipate energy proportional to VDD and are unable to recycle exhibit the same trend identified for the Buffer/NOT gate the charge.
in Section VII: within the correct operation envelope, the Fig. 1 shows the energy gain η of the LVT PFAL minimum EDP is consistently found at the lowest supply AND/NAND gate over its CMOS NAND equivalent, while voltage corresponding to the highest achievable frequency. Fig. 1 presents the energy gain η of the LVT PFAL
Both plots, additionally, provide the frequency-voltage regimes XOR/XNOR gate over its CMOS XOR equivalent. The same where gates are operational. The XOR/XNOR gate has a trend observed for the Buffer/NOT gate holds: PFAL dissipates smaller functional region than the Buffer/NOT, owing to less energy than CMOS across the entire operating region. The
its larger total capacitance (1 vs. 6 transistors per gate) energy gain has the highest magnitude for lower frequencies and the intermediate-node charge redistribution described in and for larger supply voltages. As expected, the operating Section IV. The AND/NAND gate falls between the two, con- region of the AND/NAND gate is smaller than that of the
sistent with its transistor count (8 transistors). Together with Buffer/NOT, and the XOR/XNOR is the smallest of the Fig. 12, these heatmaps complete the EDP characterisation of three, owing to their larger parasitic capacitances and the the gate library. intermediate-node charging discussed in Section IV.
3.6 3.4 4.5 3.2 4 3
Energy gain 2
3.5 2.8 1 1 2.6 2.4 2.5 2.2 2 2
1.5 1.8 10-1 10-1 0.3 0.4 0.5 0.6 0.7 0.8 0.9 0.3 0.4 0.5 0.6 0.7 0.8 0.9
VCLK [V] VCLK [V]
Fig. 15. Energy gain per period η map of the LVT PFAL Buffer/NOT gate in 1 nm with an ideal trapezoidal power-clock. η > 1 indicates PFAL energy Fig. 17. Energy gain per combination set η map of the LVT PFAL advantage. The red point marks the maximum gain η ≈ 5.2 at VCLK = 1 V, XOR/XNOR gate in 1 nm with an ideal trapezoidal power-clock. η > 1 fCLK = 1 MHz. indicates PFAL energy advantage. The red point marks the maximum gain
η ≈ 3.6 at VCLK = 1 V, fCLK = 1 MHz. 3.5 4-bit comparator 4.9 Trapezoidal
Energy gain 2
4.1 1 2.5 Sinusoidal 3.4
Energy gain 2
10-1 1 1 0.3 0.4 0.5 0.6 0.7 0.8 0.9 2.4
AND/NAND gate in 1 nm with an ideal trapezoidal power-clock. η > 1
indicates PFAL energy advantage. The red point marks the maximum gain 1.7 η ≈ 3.7 at VCLK = 1 V, fCLK = 1 MHz. Sinusoidal
1.4 Triangular Fig. 1 extends the analysis to the combinational circuits. 1.2
The energy gain of the 2×2 multiplier and 4-bit comparator
1 1 is plotted as a function of frequency for VCLK = 9 mV under three power-clock waveforms. Both circuits achieve Frequency [Hz] η > 1 across the tested frequency range under all three wave-
Fig. 18. Energy gain η of the 2×2 multiplier and 4-bit comparator realised
form shapes, confirming that the gate-level energy advantage in 1 nm process supplied with VCLK = 9 mV for trapezoidal, sinusoidal, propagates to multi-gate systems. For the 4-bit comparator and triangular power-clock waveforms. The red dots mark the maximal gain circuit, the largest gain η ≈ 2.1 is found for triangular power- for each circuit. η > 1 indicates PFAL energy advantage.
clock at fCLK = 5 MHz, while for the 2×2 multiplier, the largest gain η ≈ 4.9 is found for sinusoidal power-clock at fCLK = 1 MHz. The gain difference between the two to the ideal trapezoid, and remains within 2 % of the ideal discussed circuits results from larger losses in more complex sinusoid. At 9 mV, it is marginally more efficient than the logic trees of multi-input gates used in the comparator circuit. ideal sinusoid, attributable to the 5–8 % amplitude reduction
The CMOS reference energy for these circuits was estimated (Section VI). The P-QVCO therefore does not significantly by summing the individually simulated CMOS gate energies, degrade adiabatic energy efficiency. as described in Section VII. The reported gain therefore is an To broaden the assessment of the influence of a real power approximation. -clock on the energy recovery and to complete the gate library,
the AND/NAND and XOR/XNOR gates were driven with the designed P-QVCO and their energy dissipation was measured. C. Real Power-Clock: Influence on Energy and Driving Limits Table IV and Table V summarise the results at f ≈ 3 GHz The PFAL Buffer/NOT gate was driven by the P-QVCO for two supply voltages VDD . The AND/NAND gate follows of Section VI and its energy compared against the ideal the same trend as the Buffer/NOT: the real sinusoidal supply
trapezoidal and sinusoidal references. Table III summarises reduces energy by 26–2 % relative to the ideal trapezoid and the results at f ≈ 3 GHz for two supply voltages. The real remains within 6–9 % of the ideal sinusoid. The XOR/XNOR sinusoidal supply reduces gate energy by 16–1 % relative gate shows a slightly different behaviour: at 9 mV the real
TABLE III TABLE VI
E NERGY RATIO OF 1 NM LVT PFAL B UFFER /NOT GATE BETWEEN : P-QVCO DRIVING LIMITS FOR Rsegment = 4 MΩ, REPORTED PER TRAPEZOID , IDEAL AND SINUSOID , REAL (2 ND COLUMN ), GENERATED PHASE W. R . T. THE DIRECT POWER - CLOCK OUTPUT SINUSOID SINUSOID , IDEAL AND SINUSOID , REAL (3 RD COLUMN ) POWER - CLOCKS . WITH VCLK = 9 M V, fCLK ≈ 3 GH Z .
Eideal,trap Eideal,sin Phase shift 5◦ 5.5◦ 18.5◦ 20.2◦
f = 3 GHz, VDD = 8 mV 0.8 1.0 Amplitude loss 4.5 mV 5 mV 36.5 mV 37.5 mV f = 3 GHz, VDD = 9 mV 0.8 0.98
TABLE IV
E NERGY RATIO OF 1 NM LVT PFAL NAND/AND GATE BETWEEN : was implemented, together with a 2×2 multiplier and a 4- TRAPEZOID , IDEAL AND SINUSOID , REAL (2 ND COLUMN ), bit comparator. Three non-adiabatic loss mechanisms were SINUSOID , IDEAL AND SINUSOID , REAL (3 RD COLUMN ) POWER - CLOCKS . identified. Two specific to the PMOS/NMOS latch: threshold- Ereal,sin Ereal,sin voltage loss and redundant output-node charging. The latter is,
Eideal,trap Eideal,sin
f = 3 GHz, VDD = 8 mV 0.7 0.9 literature. One related to the complexity of the logic trees f = 3 GHz, VDD = 9 mV 0.7 0.9 of PFAL circuits – intermediate node charge redistribution.
The frequency–voltage operating boundaries of each gate were
mapped through EDP heatmaps, with the LVT Buffer/NOT supply consumes approximately 6 % more energy than the achieving a minimum EDP of 1.2 × 10−2 J·s at VCLK = ideal sinusoid. Those small differences may come from the 0.6 V and fCLK = 7.9 GHz. Across the functional region, fact that this is an approximate mathematical model of the real PFAL maintains an energy advantage over static CMOS of system implemented and simulated in Cadence Virtuoso with up to approximately 5× at the most favorable operating
ordinary differential equations. Additionally, the oscillator’s points (VCLK = 1 V and fCLK = 1 MHz). Simulations peak-to-peak output voltages are 5–8 % smaller than supply of implemented combinational circuits showed propagation VDD (Section VI). Nevertheless, these results further confirm of energy recycling capabilities to multi-gate structures. A that the P-QVCO does not significantly degrade adiabatic en- P-QVCO was designed to serve as a realistic power-clock
ergy efficiency and that a sinusoidal power-clock is a practical source. Under this real supply, the Buffer/NOT gate energy replacement for the idealised waveforms. remains within 2 % of the ideal sinusoidal reference and below The driving capability of the P-QVCO is characterised by the ideal trapezoidal energy at all measured operating points, loading each oscillator phase with progressively more PFAL confirming that energy recovery persists for real supplies.
Buffer/NOT gate instances arranged in a one-dimensional Load characterisation of the P-QVCO under a representative pipeline, with a series interconnect resistance Rseg = 4 mΩ interconnect resistance shows the phase shift and amplitude inserted between each consecutive gate, and measuring the attenuation increasing with fan-out, providing a quantitative resulting phase shift and amplitude loss relative to the un- basis for adiabatic-pipeline design.
loaded oscillator output. The 4 mΩ value was adopted from the layout-extracted 6 nm model of as a baseline. The Several limitations bounded the scope of these findings. The
1 nm process was not laid out in this work. Table VI reports P-QVCO was designed around an off-chip inductor, as the
the results for m = 0.474. Up to 2 gates per phase, the 1 nm PDK contains no inductor model. On-chip integration amplitude loss remains below 5 mV and the phase shift below would require either a process with a sufficiently high quality 5.5◦ . However, at 3 and 5 gates, the phase shift exceeds factor or alternative topologies. The CMOS reference energy 18◦ and the amplitude drops by approximately 3 mV. The for the combinational circuits was estimated by summing the
oscillator frequency falls to approximately 2.8 GHz for 5 individually simulated gate energies. Full transient simulation gate load. This RC pipeline model provides an estimate of of the equivalent CMOS circuits would tighten the comparison. the driving capabilities of the designed Power Clock. Additionally, no layout-extracted parasitics were incorporated into any of the simulations. Natural extensions include a full
VIII. C ONCLUSION
tracted parasitics. Further oscillator work could investigate the This work presented a systematic characterization of PFAL driving capabilities under physically routed interconnect rather in the TSMC 1 nm FinFET process. A small gate library than the model adopted here, optimize power consumption, such as Buffer/NOT, AND/NAND/OR/NOR, XOR/XNOR and explore alternative tank topologies in which the gate capacitance itself replaces the dedicated tank capacitor, elim-
inating the need for Cosc . Integration of an on-chip inductor
TABLE V
E NERGY RATIO OF 1 NM LVT PFAL XOR/XNOR GATE BETWEEN : presents an additional challenge. Taken together, these results TRAPEZOID , IDEAL AND SINUSOID , REAL (2 ND COLUMN ), confirm the stated hypothesis and demonstrate that PFAL SINUSOID , IDEAL AND SINUSOID , REAL (3 RD COLUMN ) POWER - CLOCKS . remains a viable energy-recovery logic family at FinFET nodes
Eideal,trap Eideal,sin
f = 3 GHz, VDD = 8 mV 0.9 1.0 the further investigation of adiabatic architectures for low- f = 3 GHz, VDD = 9 mV 1.0 1.0 energy computing.
J. M. Rabaey, A. Chandrakasan, and B. Nikolić, Digital Integrated
Circuits: A Design Perspective, 2nd ed. Upper Saddle River, NJ: Device, quantity W [nm] L [nm] Prentice Hall, 2003.
NMOS (coupling), ×4 6 2
principles,” IEEE Trans. VLSI Syst., vol. 2, no. 4, pp. 398–407, Dec. 1994.
P. Teichmann, Adiabatic Logic: Future Trend and System Level Perspec-
tive. Dordrecht, The Netherlands: Springer, 2011. captures parasitic effects more accurately than a lumped RC A. Vetuli, S. D. Pascoli, and L. M. Reyneri, “Positive feedback in model. adiabatic logic,” Electronics Letters, vol. 32, no. 20, pp. 1867–1869, Sep. 1996.
Landsiedel, “Reduction of the energy consumption in adiabatic gates by
optimal transistor sizing,” in Proceedings of the International Workshop The cross-coupled NMOS and PMOS transistors must sup- on Power and Timing Modeling, Optimization and Simulation, ser. ply a negative resistance sufficient to overcome the tank losses Lecture Notes in Computer Science, vol. 2799. Springer, 2003, pp. 309–318. and sustain oscillation. The start-up condition requires a total
N. Jeanniot, “Conception et optimisation d’une alimentation-horloge effective transconductance
et d’un réseau de distribution pour la logique adiabatique,” Ph.D. dissertation, Université de Montpellier, Montpellier, France, Nov. 2018.
A. Yousuf and K. K. M. Salih, “Design of vedic multiplier using Rp,L
adiabatic logic,” in Proceedings of the IEEE International Conference on Electrical, Computer and Communication Technologies, 2015. A factor-of-two margin was applied (Gm ≥ 2/Rp,L ≈ M. S. Alam and S. Ghimiray, “Performance analysis of a 4-bit compara- 2.6 mS) to ensure headroom. The transconductance is split tor circuit using different adiabatic logics,” in Proceedings of the IEEE International Conference on Current Trends in Computer, Electrical, equally between the NMOS and PMOS devices (gm,n =
Electronics and Communication, 2017. gm,p = Gm /2). The gm /Id methodology was used to deter- P. Andreani, A. Bonfanti, L. Romano, and C. Samori, “Analysis and mine the transistor widths . Technology-specific gm /Id , design of a 1.8-GHz CMOS LC quadrature VCO,” IEEE J. Solid-State Circuits, vol. 37, no. 12, pp. 1737–1747, Dec. 2002. gds /Id , Id /W , and C/W curves were extracted from the P. G. A. Jespers and B. Murmann, Systematic Design of Analog CMOS TSMC 1 nm PDK at Vds = 0.2 V, a bias point modelling
Circuits: Using Pre-Computed Lookup Tables. Cambridge, U.K.: the lower-swing of the oscillating voltage magnitude, with a Cambridge University Press, 2017. J. P. Duarte et al., “BSIM-CMG: Standard FinFET compact model for channel length L = 2 nm. A gm /Id operating point of 1 V−1 advanced circuit design,” in ESSCIRC Conference 20 - 41st European was selected to minimise the oscillator power consumption. Solid-State Circuits Conference, 2015, pp. 196–201. The required drain current and width for each device are
P OWER -C LOCK D ESIGN P ROCESS
Id,req This appendix documents the design procedure of the P- Wreq = . (13)
Table VII lists the transistor dimensions obtained from the
A. Design Targets and Passive Components gm /Id sizing procedure. A width sweep was performed in
The oscillator was designed for a target frequency fosc =
sustain the oscillations. The analytically calculated values
3 GHz and a supply voltage of 9 mV, targeting a single-
provided the best performance and were kept. ended output swing of at least 8 mV to provide sufficient amplitude for the combinational circuits characterised in Sec- A PPE
FAQ
Cadence Lab — Bangalore
Simulation, control and hardware support for final-year robotics projects.
Stacks
Worlds
Digital Twin
Control
Robots
Offline
Bring-up