Enquire Now
70+ Topics · Cadence · Siemens EDA · Synopsys · Keysight ADS · LTspice · Bangalore 2026

Peak Detector Schematic and Simulation Cadence

Schematic · Layout · Simulation · Signoff — Final-year Cadence / EDA topics with Virtuoso, Spectre, Innovus, Siemens EDA, Synopsys, Keysight ADS and LTspice. Notes, report, PPT and viva support from Bangalore.

72+
Related Topics
7
EDA Tools
4.9★
573 Ratings
Schematic Simulation Layout Physical Design Performance Capstone

USA India Barbara

sarab@cs.utah.edu neelam.surana@alumni.iitgn.ac.in USA pranjali.jain@alumni.iitgn.ac.in

ABSTRACT 1 INTRODUCTION

In this paper, we propose a “full-stack” solution to designing high With the increasing number of cores on-chip , additional capacity and low latency on-chip cache hierarchies by starting at memory is needed to feed these cores. Emerging workloads have the circuit level of the hardware design stack. First, we propose a become significantly memory intensive and have large working set novel Gain Cell (GC) design using FDSOI. The GC has several sizes [54, 73]. These factors have necessitated the need for high

desirable characteristics, including ~50% higher storage density capacity, low latency, on-chip caches. and ~50% lower dynamic energy as compared to the traditional Several research efforts have been made to increase capacity and 6T SRAM, even after accounting for peripheral circuit overheads. contain latency of on-chip caches, which has led to ever-increasing We also exploit back-gate bias to increase retention time to 1.1 cache capacities and deeper cache hierarchies . However, the

ms (~60× of eDRAM) which, combined with optimizations like memory technology, which makes up the bulk of on-chip caches, staggered refresh, makes it an ideal candidate to architect all levels has remained unchanged. On-chip caches, at almost all levels of of on-chip caches. We show that compared to 6T SRAM, for a the memory hierarchy, have been devised using 6T SRAM. Even given area budget, GC based caches, on average, provide 29% and though SRAM suffers from low areal density, high leakage power

36% increase in IPC for single- and multi-programmed workloads, and high dynamic energy requirements compared to other memory respectively on contemporary workloads including SPEC CPU 2017. technologies [16, 68, 85], the latency superiority of SRAM, and its We also observe dynamic energy savings of 42% and 34% for single- compatibility with logic fabrication technology has made it indis- and multi-programmed workloads, respectively. pensable for creating low-latency caches.

We utilize the inherent properties of the proposed GC, including Recently, a number of alternative memory technologies have decoupled read and write bitlines to devise optimizations to save started to emerge, and have been evaluated for use in caches. Con- precharge energy and architect GC caches with better energy and tenders for SRAM replacement include non-volatile memory tech- performance characteristics. Finally, in a quest to utilize the best nologies like Spin-Transfer Torque RAM (STT-RAM) [75, 77, 88]

of all worlds, we combine GC with STT-RAM to create hybrid and Phase Change Memory (PCM) , as well as volatile ones like hierarchies. We show that a hybrid hierarchy with GC caches at embedded DRAM (eDRAM) . L and L2, and an LLC split between GC and STT-RAM, with In addition to providing data non-volatility, both STT-RAM and asymmetric write optimization enabled, is able to provide a 54% PCM provide 3-4× density benefits over 6T SRAM [56, 77], making

benefit in energy-delay product (EDP) as compared to an all-SRAM them attractive candidates for high capacity caches. However, there design, and 13% as compared to an all-GC cache hierarchy, averaged are drawbacks inherent to both technologies, including higher write across multi-programmed workloads. energy (up to ~5× that of SRAM) and access latencies (~1.5× read, 5× write latency for STT-RAM, PCM is worse) [56, 77]. In most

KEYWORDS cases, the drawbacks outweigh the benefits, rendering these technolo- Cache Memory, Emerging Memories, Gain Cell gies suitable for use only in last-level caches, where both capacities and access latencies are expected to be higher. eDRAM has also been evaluated as a candidate for architecting caches [16, 86]. It has * Both authors contributed equally to this research. found adoption in multiple recent products, including IBM’s Power

† This work was carried out while author was at Ashoka University. ‡ series , Intel’s Haswell and Microsoft’s Xbox 3 , This work was carried out while authors were at Indian Institute of Technology, Gandhinagar. again as a technology for LLCs. Since eDRAM is a DRAM variant, it provides higher density compared to SRAM and has favorable

Data Storage Latch Magnetization Phase Capacitor MOS MOS MOS MOS

Read/Write Time Short/Short Short/Long Short/Long Short/Short Short/Short Short/Short Short/Long Short/Short Read/Write Energy X/Y X/5Y X/5Y 0.5X/0.5Y 0.5X/0.5Y 0.5X/0.5Y 0.5X/0.5Y 0.5X/0.5Y

Leakage/Yield High/High Low/High Low/High Low/Low Low/High Low/High Low/High Low/High

Retention Time - - - 2 us 2 us 2 us 1.6 ms 1.1 ms

Decoupled Bitline

the presence of a suitable SRAM replacement, many of their pros can be combined to architect high capacity caches that have similar latency profiles as that of an SRAM based hierarchy. In this paper, we attempt such a design, starting from ground up. First, we propose a novel FDSOI MOSFET based 2T Gain Cell, which provides high storage density, low access latencies, and a high DRT. This makes the cell amenable for use at all levels of the cache hierarchy.

Figure 1: eDRAM LLC energy breakdown. to 6T SRAM, the proposed GC array offers a 50% reduction in

read/write energies, 50% reduction in the area while keeping access latency unchanged. It exhibits low leakage energy and has modest access latency profiles as compared to NVMs. However, the tra- refresh requirements. A comparison of advantages of the proposed ditional 1T1C eDRAM cells suffer from low Data Retention Times GC over competing memory technologies for caches is presented in (DRTs) of 2 - 5 µs , requiring frequent refreshes in many Table 1. The main contributions of this work are summarized below:

eDRAM based design. These refreshes cause significant energy con- sumption as the data from an eDRAM row has to be read and written back [16, 86], making addressing refresh operations as the primary • We propose a novel, FDSOI based 2T Gain Cell, specifically challenge in designing eDRAM caches. The energy consumption for use in on-chip caches, and exploit its back-gate bias fea- breakdown of dynamic and refresh energies for single programmed ture to reduce leakage power by 99% and increase retention

SPEC CPU20 and PARSEC workloads in an eDRAM LLC (simu- time by ~60×, compared to conventional GC and eDRAM, lation parameters are listed in Section 7) is shown in Figure 1. As reducing the need for frequent refreshes. We combine this al- can be observed, refresh energy is many times higher than dynamic ready large refresh window with optimizations like staggered energy, which makes optimizing refresh operations an essential con- refresh, to design practically refresh-free GC caches. As a

sideration for designing eDRAM based caches. result, proposed GCs can be used at all levels of the cache Apart from energy overheads, an eDRAM row, and hence a por- hierarchy. tion of the cache, is unavailable for the duration of the refresh period, • We show that GC based sub-arrays exhibit 2× area advantage leading to performance overheads. To counter the shortcomings of and similar latency characteristics as compared to SRAM,

traditional eDRAM, another eDRAM variant, Gain Cell (GC) has even after accounting for overheads of peripheral circuits. been proposed [30, 76], which has multiple advantages over 1T1C This helps architect higher capacity caches within the same eDRAM. These include a cheaper, logic-compatible fabrication pro- area and latency budgets. As a result, for single-programmed cess and the absence of a dedicated capacitor per cell; GCs use workloads, iso-area caches architected using proposed GCs at

the transistor’s parasitic capacitance for data storage. Finally, GCs all levels of the hierarchy exhibit a 42% reduction in dynamic provide non-destructive reads and decoupled read/write bitlines, energy and a 29% increase in IPC as compared to SRAM. resulting in lower access latency and energy consumption . Multi-programmed workloads exhibit similar behavior. We Despite these advantages, traditional 2T and 3T GCs have not further show that the proposed GC scales well at smaller

found adoption widespread since they suffer from low DRTs, hence technology nodes, and retains its advantages over SRAM. requiring frequent refreshes. While 4T GCs with higher DRTs have • We utilize the inherent decoupling in the read and write bit- been proposed , they suffer from higher write energies and lower lines in proposed GCs to save precharge energy between density, making them unattractive as 1T1C eDRAM replacements. consecutive writes, and clubbing up to 70% writes with reads,

In any case, the presence of refresh makes existing GCs useless at thereby improving performance and reducing dynamic en- any level of cache, other than LLCs . ergy consumption by 13%, as compared to GC based caches As a result, even though each one of these technologies has its lacking this optimization. We also explore optimizations like pros and cons, they cannot be used as a drop-in replacement for no-refresh policy, where a line is invalidated if it is not ac-

SRAM caches, especially for levels closer to the CPU. However, in cessed during its DRT period, and show that this can be used

HyGain: High Performance, Energy-Efficient Hybrid Gain Cell based Cache Hierarchy

proposals still suffer from low DRT leading to enormous refresh energy, as shown in Table 1. Earlier, FDSOI devices have been fabri- cated in 1 nm and are expected to scale down, considering their advantages of higher DRT over FinFET[19, 38]. We implement an n-type 2T GC on an FDSOI device [1, 65] and exploit its back-gate bias feature to lower the leakage current, thus improving the DRT.

Data retention time (DRT) of a transistor is directly linked with the

Figure 2: The effect of back-gate bias voltage on (a) Threshold leakage current of the transistor in the OFF condition. The leak- Voltage (VT H ) (b) Leakage Current (IOFF ) and On-state (ION ) age current (IOFF ), in turn, exponentially depends on the threshold in n-type FDSOI transistor. voltage (VT H ) of the transistor. In the recent past, with technol- ogy scaling, leakage has prohibitively increased in bulk-MOSFETs

necessitating significant manufacturing efforts and additional fab- rication steps to control the leakage current. Silicon-on-Insulator (SOI) technology offers a promising solution to deal with leakage current due to an additional handle to control the threshold volt- age using its back-gate biasing. The junction leakage currents are significantly reduced in a fully-depleted SOI (FDSOI) as an oxide layer removes the p-n junction from the substrate, and has been used in a few commercial offerings, including IBM’s POWER 8 .

Due to this oxide layer, the substrate works as a second gate or the

back-gate, and can be biased to change the threshold voltage (a) Proposed 2T GC (b) 6T SRAM Cell. of the FDSOI transistor which, in-turn, exponentially reduce the leakage current [13, 24, 61, 81, 89, 90], and increases the DRT of

We have implemented the n-type transistor using ST Microelec-

tronics 28-nm FDSOI technology and simulated it using Cadence as an effective mechanism to eliminate refreshes in caches Virtuoso. Figure 2(a) captures the effect of back-gate biasing (Vbg ) closer to the CPU without performance penalties. on threshold voltage, when Vbg is varied between −2V to 2V . Fig- • Finally, to create high capacity, low latency caches, we evalu- ure 2(b) shows the exponential dependence of IOFF on Vbg , and as

ate several hybrid cache hierarchies by incorporating emerg- Vbg becomes more negative, the leakage current exponentially drops ing technologies like STT-RAM with GCs to design caches by almost four orders of magnitude, up to 1 pA/µm. Figure 2(b) that can provide up to 4× higher capacity, compared to iso- also shows the improvement on ION (on-current) when Vbg is pos- area SRAM caches. For multi-core workloads, hybrid caches itive. Improvement in ION makes the transistor operate faster, thus

architected with GCs, combined with asymmetric writes op- reducing delay. Thus, back-gate biasing can be used both to improve timization, can provide 43% performance and 44% energy performance (latency) and reduce leakage. The schematic of the benefits as compared to iso-area SRAM caches. proposed FDSOI based Gain Cell is shown in Figure 3a. During The remainder of this paper is organized as follows. We provide the hold condition, -VDD is applied to the back-gate bias of the device,

circuit level implementation details for proposed GC in Section 2 which reduces the leakage from W1, improving DRT. and details of the architectural implementations of GC caches in

Section 3. Section 4 presents the experimental evaluation setup, 2.2 Overheads of back-gate biasing

while Section 5 analyses the energy and performance implications of Leveraging back-gate bias for DRT improvement leads to the reduc- the implementations. Section 6 proposes optimizations by exploiting tion of ON current of n-type FDSOI, as depicted in Figure 2(b), and intrinsic properties of GC, while Section 7 evaluates hybrid cache hence, higher cell access latency. To capture the best of both worlds, hierarchies. Sections 8 discusses the scalability of GC at lower we apply zero back-gate bias during read and write operations, and

technology nodes and provides an assessment of the no-refresh -VDD during the hold condition. This can be implemented with the policy. Finally, we discuss related work in Section 9 and conclude in circuit proposed in . Section 10. Figure 4 shows the layouts of SRAM and proposed 2T GC. To implement back-gate bias, a single contact can be shared across all 2 BACK-GATE BIASED GAIN CELL the cells of a row in an array (as shown in Figure 4(b)). This keeps

Gain Cells have started gaining traction owing to their logic-compatible the area overhead minimal. Since the back-gate oxide thickness fabrication process, small area footprint, and low energy require- is quite large (~2 nm), back-gate carries only 5% of the front- ments [28, 30, 76, 79]. GCs have been fabricated and tested in gate capacitance, which keeps switching power overhead less than FinFET , bulk , and FDSOI processes. However, these 5%, while increasing DRT. Additionally, since Vbg and row-signal

(WWL) switching happens in parallel, delay of back-gate is over- Table 2: Working of the Proposed Gain Cell shadowed by WWL’s delay. Hence, back-gate bias causes no latency penalty, negligible area penalty, and has only ~5% of the switching Operation WBL WWL Back-gate(bg) RBL RWL

Read − 0 0 VDD (floating) 0

power overhead. Write Data VDD 0 VDD VDD

2.3.2 Read Operation. For a read operation, Read Bitline (RBL),

which is a shared signal across the column, is first precharged to

VDD . Then, to read data, RBL is kept floating. Active low signal

to Read Wordline (RWL) is used to read the data. If data stored in Q is 1, RBL discharges; otherwise it remains at VDD , which is sensed by a sense amplifier. During the read operation, energy is consumed in the switching of RBL and RWL. Most importantly,

Figure 4: Layout of (a) 6T SRAM Cell (b) Proposed 2T Gain read operation in GC is non-destructive, since RBL is decoupled

Cell (all cells in a row share a single back-gate, layout shows 2 from the Q node . GCs with shared B-G).

2.3.3 Hold Condition. Periods where the cell is neither read nor

written to is known as the hold condition. This is important since the cell still needs to retain data during this period, unlike SRAM where data is retained due to cross-coupled inverters shown in Figure 3b.

The charge in Q node leaks from W over time, necessitating refresh

operations for data restoration, before the DRT window closes.

2.4.1 Data Retention Time (DRT). The proposed GC uses back-gate

bias voltage of -VDD during hold operation, leading to reduction in

Figure 5: 10K, Monte-Carlo waveforms of 1 and 0 decay (a) leakage current from the W transistor, improving DRT. To quantify

Traditional 2T Gain Cell (b) Proposed 2T Gain Cell. DRTs, we performed Monte Carlo simulations considering the stan- dard 6-σ local and 1-σ global process variations. Figure 5 shows the data degradation of conventional and proposed 2T GC, for 10K M-C simulations. DRT is measured at a point where the data can be read without error. Typically, VDD /3 is a sufficient margin to read data properly, and we consider the worst-case DRT as the refresh interval, making these results more pessimistic than usual. For conventional

2T GC , DRT obtained is 19µs by considering 100% yield (Fig- ure 5(a)), which is consistent with . For the proposed GC, the data decay has significantly slowed down, as seen from Figure 5(b).

Figure 6: Leakage Power of 6T SRAM, Traditional GC & Pro-

improves to 1.1 ms, which is ~60× higher than the case when no posed GC back-gate bias is applied. This is first such GC proposal built with

2 transistors leading to significant gains in DRT over existing GC

designs, except 4T GC designs where a similar DRT is achieved

Figure 3(a) shows the schematic of the proposed n-type 2T GC. GC 2.4.2 Leakage Power. Since the proposed 2T GC has a smaller

has decoupled read and write operations and has non-destructive number of leakage paths compared to SRAM (schematic in Figure 3), reads, unlike 1T1C eDRAM. The input signals for these operations it inherently has lower leakage power. Additionally, we have used are illustrated in Table 2. Working of proposed 2T GC is similar to back-gate bias to further reduce the leakage current significantly. We conventional 2T GC, except that we use back-gate bias during hold compare the leakage power of 6T SRAM, GC without back-gate

condition. bias, and proposed GC with back-gate bias in Figure 6. We show that proposed GC has ~99% reduction in leakage power as compared to

2.3.1 Write Operation. Write Bitline (WBL) is shared across an 6T SRAM.

entire column in an array and has a large capacitance. Write Wordline (WWL) runs along the row in the array. To write to the cell, data is 2.4.3 Area. Figure 4 compares the layout of the 6T SRAM cell and first transferred to WBL, and then a row-signal (WWL) is used to 2T GC at 2 nm technology. SRAM cell takes 0.6 µm , whereas transfer data to the Q node. the proposed GC takes 0.2 µm , which is 40% of the SRAM cell.

At the cell level, the proposed GC takes only 0.4× area com-

pared to the 6T SRAM cell. Layout of the SRAM cell is drawn in a very efficient way and have area efficiency of 80%-90%[40, 68, 85].

Considering this, at the cache level, GC can have ~2× capacity as

compared to SRAM cache. Even though GC has 2.5× benefits at the cell level, at the cache level, it reduces to 2.0× due to peripheral cir- cuitry overhead. In the rest of the paper, for the iso-area comparison, we have considered 2× capacity of GC compared to the SRAM.

Cache Level 32kB L 256kB L 8MB L

the mechanism by which it maps to various cache configurations. SRAM 0.475(2) 1.34(5) 2.81(10)

Latency (ns) (in cycles)

Then we compare the architectural benefits of GCs over SRAM. GC 0.42(2) 1.20(5) 2.55(9)

Read/Write Energy SRAM 0.75/1.1 2.18/3.1 7.5/11.8

GCs are arranged as sub-arrays, in a typical row-column fashion. per bit (pJ) GC 0.41/0.6 1.15/1.6 4.1/5.8 A cache can then be mapped to multiple sub-arrays, as dictated by its Write Energy/bit for Same SRAM Same as Write Energy

Bit (0->0 or 1->1) (pJ) GC 0.2 0.7 1.9

capacity. From an extensive design space exploration, we conclude Leakage/bit (pW) (SRAM/GC) 13.27/0.0 that the sweet spot for minimum latency and peripheral circuitry Refresh Interval(ms)/ Period per line(ns) 1.12/1.5

Refresh Energy/bit (pJ) 1.8

overheads lie at a sub-array size of 256×5 bits, or 1 KB. Hence,

GC caches can be architected such that each way, across all cache

sets, maps to one sub-array. As a result, looking up a cacheline (64B) signals while performing refresh (no extra delay penalty). For stag- is the same as looking up a row of this sub-array, as shown in the gering the refresh across different rows, the counter times out after Set0-Way to Row mappings in Figure 7. In cases where combined every DRT/N time (N is number of rows in subarray). size of a way is >1 KB, we keep adding sub-arrays, until all the Next, we compare and contrast the architecture level character-

sets have been accounted for. For example, the right hand side of istics of on-chip caches devised using SRAM and GCs. We ex- Figure 7 depicts a cache where one way is mapped to two 256×5 tract SRAM and GC energy and latency parameters using an en- sub-arrays. hanced CACTI model and present these results in Table 3. However, this prohibits mapping of any cache configuration where These results were also validated using SPICE simulations using ST-

the combined capacity for one way is smaller than 1 KB, as illus- microelectronics 28-nm FDSOI CMOS technology. We observe that trated in the left half of Figure 7. For these caches, we keep the for every cache level, the dynamic read and write energies for a GC design choice of mapping an entire way to one sub-array, while cache are reduced by at least 50%, as compared to an SRAM one. reducing the size of the sub-array. For example, in the case of a This is because the proposed GC requires just one bitline per read or

64KB, 16-way cache, an entire way (4KB) is mapped to a 64×5 write access, thereby reducing a significant fraction of the dynamic bit sub-array. This ensures that a 6 B cacheline lookup is not spread energy consumed in switching of bitlines and word lines [52, 78]. across multiple sub-arrays.

Additionally, owing to decoupled read and write like 8-T SRAM

caches operations, GC has slightly lower access latencies as compared to

SRAM, allowing for similar cycle time access as that of a SRAM

Figure 8 shows the block diagram of the proposed GC cache. cache, for a given processor frequency. Another trait of GC is high Compared to SRAM cache, GC has additional refresh counters at density, which enables us to fit a similar capacity cache in half the each level of cache. As per concurrent refresh, the refresh counter area. Alternatively, in a given area budget, we can implement a will generate signal to refresh the same row in all subarrays at the higher capacity cache by increasing associativity. As shown in Fig-

same time. BGB signal is generated along with RWL and WWL ure 9, iso-area access latencies of GC based caches at all levels of

the hierarchy are similar to SRAM caches, with additional benefit energy and performance implications in Sections 5.1 and 5.2, re- of GC caches having twice the capacity. Doubling the capacity of spectively. These workloads, listed in Table 5, are simulated for 2 an SRAM based cache increases the area by 2.0×. Not only that, it billion instructions each, after a 5 million warmup period. Addi- also increases access latency of caches by at least 30%. As a result, tionally, to include variations in the application behavior and test

GC based caches allow for twice the capacity in the same latency against multi-programmed workloads, we divide the workloads in six and area budget, for every level of cache. sets, each consisting of 8 benchmarks, which represents: memory- intensive applications (MEM_HIGH), applications with average

3.1 GC Based Cache Proposals memory access (MEM_MED), applications with sparse memory

Using these observations, we propose the use of GCs at various access (MEM_LOW), and three random mixes of 8-workloads levels of caches, from all on-chip caches architected using GCs (mix1-3). This classification is done based on memory accesses (ALL-GC) to just last-level cache (LLC-GC) or L cache (L1-GC) per kilo instructions to caches - higher accesses means more mem- being GC, and compare with the baseline case where all caches ory intensive. Also, we create six homogeneous multi-programmed

are implemented with SRAM (ALL-SRAM). Since GCs provide workloads (8x*) - 8 copies of the same benchmark, each per core. excellent density benefits over SRAM, we examine the energy and

Table 4: System Configuration (ALL-SRAM)

performance implications of GC over SRAM for both iso (cache) capacity and iso-area. For iso-capacity (-CAP), SRAM and GC

Processor 8-core, 3.4GHz, x86_ ISA, 19-stage OOO

caches are compared with the same cache size, which indicates lower Decode, Rename, Fetch Width 4-7 fused, 4, 6 instructions per cycle on-chip area usage by GC caches. While in the case of iso-area (- Issue, Dispatch, Commit width 4, 6, 4 fused µ-ops per cycle

ROB/Branch misprediction 1 entries/8 cycles penalty

AREA), the GC caches are doubled in capacity by increasing their L1-I/L1-D cache 3 KB, 8-way & 2 cycles. 64B line associativity, while retaining latency characteristics. We maintain L cache 2 KB, 8-way & 5 cycles. 64B line

L cache Shared 8 MB, 16-way & 1 cycles. 64B line

the tag array in SRAM; only the data arrays are replaced with GCs. 4096MB DDR3, 1 ns access,

3.2 Handling Refresh in Gain Cell Caches

One of the biggest challenges in GC caches is the need to refresh. In Table 5: Workloads addition to adding energy overheads, the cache is made unavailable for access during refresh operations, which adversely affects perfor- Multi-Programmed mance. We use a staggered, concurrent refresh mechanism to Single-Programmed MEM_HIGH MEM_MED

reduce unavailability of GC cache. SPEC CPU20 cactuBSSN_r, mcf_r, xalancbmk_r, ferret, perlbench_r, gcc_r streamcluster, gcc_r, cam4_r, bwaves_r, As explained earlier in this section, one way of the cache is bwaves_r, mcf_r canneal, omnetpp_r, deepsjeng_r, wrf_r, mapped to a row in the GC sub-array. Refreshing one row of the cactuBSSN_r, parest_r facesim, perlbench_r povray_r, freqmine

povray_r, lbm_r MEM_LOW mix sub-array takes 3 ns. We refresh one row in a sub-array at a time and omnetpp_r, wrf_r raytrace, parest_r, bwaves_r, xz_r, xalancbmk_r, cam4_r fluidanimate, nab_r, wrf_r, raytrace, iterate over all the rows in a round-robin fashion in the course of the deepsjeng_r, imagick_r dedup, imagick_r, roms_r, dedup, 1.1 ms refresh window, which is the DRT of an individual cell. nab_r, roms_r, xz_r roms_r, lbm_r lbm_r, freqmine

PARSEC mix mix

A refresh is done by reading the sub-array in the first 1.5 ns of the canneal, dedup perlbench_r, ferret, mcf_r, cactuBSSN_r, refresh window. Data is written back to the row in the second half of facesim, ferret parest_r, canneal, povray_r, xalancbmk_r fluidanimate, freqmine omnetpp_r, cam4_r, deepsjeng_r, imagick_r the window. As a result, the sub-array is available for write in the raytrace, streamcluster nab_r, streamcluster facesim, fluidanimate

first half and a read in the second half. This is made possible due to 8x* - Running 8 copies of the same benchmark the presence of separate read and write bitlines. This optimization increases the availability of the sub-array and hence, the associated way – it is now unavailable only for 1.5 ns every 4.3 µs. Since 5 EVALUATION OF GC CACHES the tags are maintained in SRAM, the cache can still be accessed to In this section, we quantify the benefits of various GC based architec-

check for hits/misses. tures and compare them with SRAM based caches. The evaluation We carry out a detailed analysis of performance and energy includes a study of memory subsystem energy, performance, and the implications of this refresh policy, in Section 5.3, and conclude impact of refresh operations. that the performance overheads are minimal, since a tiny fraction (0.003%) of cache accesses happen concurrently with refresh, lead- 5.1 Energy Analysis

ing to a worst-case 1.7% reduction in performance, as compared to While most emerging memory technologies exacerbate energy re- the SRAM baseline. quirements of the memory subsystem [16, 49, 91, 92], GCs, on the contrary, provide significant energy savings over SRAM. In com-

4 EVALUATION METHODOLOGY parison with the baseline SRAM, at the array level, proposed GC

We evaluate our proposed architectures by using an 8-core system design consumes ~46% less energy per read and 40-50% less energy with configuration listed in Table 4, simulated using Sniper . This per write, as shown in Table 3. configuration is used as the baseline for evaluation (ALL-SRAM). The ALL-SRAM case, where all levels of caches are assumed For iso-area GC caches (-AREA), cache capacity is doubled by dou- to be SRAM, is used as the baseline. We first study ALL-GC-CAP,

bling associativity. We test our proposals against 2 benchmarks where we replace all SRAM caches with the same capacity GC from the SPEC CPU20 and PARSEC suites and study caches (iso-capacity). Next, we evaluate iso-area GC caches with

Figure 10: Dynamic energy (Cache + Main Memory) for single & multi-programmed workloads. Normalized against ALL-SRAM.

double capacity. For this, we evaluate three configurations: (a) L1- 5.2 System Performance GC-AREA, where we replace L cache in ALL-SRAM with a dou- Figure 1 illustrates the system performance in terms of instruc- ble capacity GC cache. (b) LLC-GC-AREA, where we replace last- tions per cycle (IPC) for various proposals, using single and multi- level cache in ALL-SRAM with double capacity GC, and (c) ALL- programmed workloads. We observe that the performance differ-

GC-AREA, where we replace all levels of caches in ALL-SRAM ence between iso-capacity SRAM caches (ALL-SRAM) and GC with double capacity, iso-area GCs. caches (ALL-GC-CAP) is negligible - 0.1% drop in IPC for GC Figure 1 compares the dynamic energy consumed by the mem- caches on average, with respect to ALL-SRAM. All GC based iso- ory subsystem in proposed architectures, for both single and multi- area caches (ALL-GC-AREA) exhibit performance gains as com-

programmed workloads. We calculate the dynamic energy consump- pared to the baseline. Average performance increase of 29% and tion of caches and main memory by taking the product of the total 36% is observed across single and multi-programmed workloads, number of accesses to each level with the energy consumption of per respectively. In the case of multi-programmed workloads, memory- access using the per bit cache access energy (mentioned in Table 3). intensive workloads tend to benefit most. We observe a 27% per-

Access energies of main memory are obtained from and are formance increase for MEM_HIGH workloads mix, as compared presented in Table 4. We observe that any cache hierarchy devised to the baseline. We verified our results on an aggressive processor using GCs exhibits savings in dynamic energy. In L1-GC-AREA, configuration, with better prefetcher, replacement policy and DRAM where only the L is architected using proposed GCs, results in a access latency [35, 63, 74], and observed similar benefits (<4% IPC

36% reduction in dynamic energy, averaged across all the single pro- drop from above reported benefits). grammed benchmarks. The LLC-GC-AREA configuration, which Applications like streamcluster, canneal, and mcf_r have working replaces the SRAM LLC with a double capacity GC cache, also ex- set sizes that exceed the capacity of baseline SRAM caches. Hence, hibits a 4% average reduction in dynamic energy. Similar results are when the LLC size is doubled (LLC-GC-AREA), they see large

obtained for multi-programmed workloads as well. For the memory- performance improvements (>200%) as working sets can reside intensive mix (MEM-HIGH), the energy savings of L1-GC-AREA on caches. While, many applications like cactuBSSN_r and gcc_r and LLC-GC-AREA stand at 25% and 9%, respectively. have working sets that reside in the on-chip memory and leverage Finally, using GC for all levels of the cache increases these gains larger cache size to fit working sets in the L cache. As a result, they

tremendously. On average, across the single programmed work- achieve significant performance improvements in L1-GC-AREA loads, ALL-GC-AREA achieves 42% (34% in the case of multi- implementation (drop-in accesses to next level caches by >90%) but programmed) reduction in dynamic energy consumption as com- not in LLC-GC-AREA. On average, L1-GC-AREA and LLC-GC- pared to the ALL-SRAM baseline. Even in cases where the area AREA achieve IPC improvements of 13% and 15%, respectively.

density benefits of proposed GC are not being utilized, i.e., in the sub- optimal configurations of ALL-GC-CAP, where all SRAM caches 5.3 Impact of Refresh are replaced with equal capacity GC caches, we observe an average reduction in the dynamic energy of 34% and 28% for single and Regular GC based caches are required to refresh cells at regular multi-programmed workloads respectively. intervals. Unfortunately, this has adverse effects on both energy

Compared to the baseline, applications like streamcluster and and performance. 1T1C eDRAM and traditional 2T, and 3T GCs mcf_r, that have large number of memory accesses, achieve up to can have huge refresh energy overheads, accounting for up to 97% 80% reduction in dynamic energy for iso-area GC LLCs (LLC- of total LLC energy, as was observed in the experimental results GC-AREA). Increasing LLC capacity allows the working set of presented in Figure 1.

these applications to reside in the cache, reducing the number of off- However, for caches designed using the proposed GC, owing chip accesses substantially (by ~99%). Reduced off-chip accesses to high DRTs and staggered refresh mechanisms, we observe that reduce the high off-chip dynamic energy, resulting in massive energy the refresh energy consumption is minimal, assuming the most pes- savings. Additionally, in a large, many-core CPU running at a low simistic scenarios. For experiments carried out with an all GC based

voltage, leakage from on-chip caches contributes substantially to the cache subsystem (ALL-GC-AREA), which should exhibit the worst chip’s power draw . Proposed GC, with a large savings of 99.3% case refresh energy consumption profile, we observe that on av- in leakage energy, as depicted in Figure 6, helps reduce these costs erage, across all single programmed benchmarks, refresh energy substantially. contributes <3% (6%, at max for ferret) of the total energy con-

sumption of all caches. We illustrate these observations in Figure 12.

Figure 12: Percentage of accesses to the cache which arrive when the cacheline is being refreshed (y1-axis). Breakup of energy con-

sumption of caches, normalized to total energy (Dynamic + Refresh) (y2-axis). Configuration used is ALL-GC-AREA.

contribution to the total energy, as depicted in the histogram on present the ratio of dissimilar bit writes to total bit writes, for each y2-axis. cache. Lower ratio implies that a lot of the data being written is Besides, the refresh operations do not affect performance ad- similar to the value of the WBL, and hence can be written with versely, as demonstrated from ALL-GC-CAP results from Sec- lower write energy. Accordingly, we calculate the overall dynamic

tion 5.2. This is evidenced by the fact that caches spend only 0.008% energy consumption (cache + main memory) and compare it with of the time on refresh, on average. Our experiments show that, on baseline ALL-SRAM and ALL-GC-AREA proposals on y2-axis. average, ~0.003% of accesses to caches were made during refresh We observe that most bits – 76%, averaged across all levels of caches, interval throughout the entire simulation (Figure 12, y1-axis). in write data are similar to the write bitline’s value. SPEC CPU20

workloads rarely have writes which are larger than 8Bytes . The 6 ASYMMETRIC WRITES rest of the cacheline is re-written with the same value. This is true for all data caches, with L1D exhibiting as high as 94% similarity, To read/write a value in a 6T SRAM cell, the bitlines first have to be averaged across all benchmarks. Even in cases, where the initial precharged to VDD . Then, wordlines are turned on to access the cell access was a miss, and the existing cacheline has to be replaced with

value. On a read, both bitlines are precharged to high (VDD ) while a new one being brought in, we observe significant data similarity on a write, one bitline is driven to high and other to low [68, 85]. In between the new and the old lines. For instance, L I-cache, which proposed GC, we have separate bitlines for read (RBL) and write only experiences writes as insertions of new cachelines, observes (WBL), as depicted in Figure 3a. Due to this decoupling, there is a 55% data similarity between the old and new lines. As a result,

no need to precharge the WBL to VDD before every access, unlike we observe a 13% reduction in dynamic energy consumption, with in SRAM where the bit lines are pre-charged to VDD before every respect to ALL-GC-AREA, and 50% compared to baseline SRAM- write. To perform a write, the WBL is precharged to high or low, based cache subsystem (ALL-SRAM). Applications with higher depending on the value to be written. For instance, to write a 0, write ratios tend to save more energy, for example, parest_r,

WBL will initially be connected to ground to establish the voltage

omnetpp_r, xalancbmk_r, and povray_r. Therefore, we show that, difference to drain the charge in the cell to 0. Thus, if WBL was because of its inherent structure, GC can take advantage of write already set to 0, driving it again to 0 would require no energy. Such similarity in data to further save on dynamic energy. cases arise when the consecutive writes by WBL are the same, i.e.,

Additionally, due to decoupled bitlines, reads and writes to the

0 → 0 or 1 → 1 transitions. In these cases, the second write will have same sub-array can be done in parallel. We use this property to a 0.4 − 0.67× lower energy consumption as compared to a similar overlap writes with simultaneously occurring reads to the same sub- transition in 6T SRAM, as depicted in Table 3. array. We note that 40% of all writes could be overlapped with some

We quantify the overall dynamic energy savings due to such asym-

reads, represented by line-graph Overlaps in Figure 14. Most of the metric writes by calculating the number of dissimilar bits between overlaps happen in L1-D cache, as L and L caches have more two consecutive writes. Dissimilar bit writes are serviced normally, sub-arrays and experience a smaller number of writes than L1. By while writes with similar bits are serviced with reduced energy. The hiding the latency of these writes, we observe a ~2% increase in reduced energy parameters for caches, extracted from CACTI, are de-

IPC compared to ALL-GC-AREA, presented by GainbyOverlaps

picted in Table 3. We perform this experiment with ALL-GC-AREA in Figure 14. To further take advantage of decoupled bitlines, smart configuration and present our results in Figure 13. On y1-axis, we

Figure 13: Dissimilar bits among consecutive writes, per cache level (y1-axis, %). Dynamic Energy (Cache + Memory), normalized to

helps it achieve 2.3% (4% for multi-core) improvement in IPC over ALL-GC-AREA, which makes a case for an eDRAM-based LLC.

However, eDRAM has large refresh overheads. Even the refresh-

optimized eDRAM (GC-GC-eDRAM) results in 21% (12% for multi-core) more total energy consumption than ALL-GC-AREA.

Traditional eDRAM results in much worse energy overheads – 6.8×

higher energy consumption as compared to ALL-GC-AREA.

STT-RAM, with the same 2× density benefits over GC, suffers

from much longer access latencies, which have been enumerated Figure 14: Percentage of writes that overlap with reads (y1- in Table 1. As a result, configurations with STT-RAM LLC experi- axis). Accordingly, the increase in IPC over ALL-GC-AREA enced a performance drop of 7% with respect to ALL-GC-AREA, (y2-axis). even though the number of off-chip requests actually dropped signifi-

cantly by 13%. However, in realistic cases (multi-core runs) that take

Table 6: LLC parameters for different technologies advantage of larger LLC, GC-GC-STTRAM performs slightly better

than ALL-GC-AREA. Consequently, with smaller off-chip accesses, eDRAM STTRAM Hybrid (8MB GC energy consumption reduces by 5%, compared to ALL-GC-AREA,

32MB 32MB +16MB STTRAM)

Read Latency (ns) (cycles) 5.1 (18) 2 (89) 2.9 (10), 2 (89) concluding that STT-RAM LLC would be a better design. Write Latency (ns) (cycles) 5.1 (18) 6 (204) 2.9 (10), 6 (204) In an effort to get the best of all worlds: utilize higher density of

Read/Write Energy / bit (pJ) 5.2/6.1 5.35/7.8 3.81/5.52, 5.35/7.8

Refresh Interval/Period 0.02ms/4ns -/- 1.12ms/1.5ns, -/- STT-RAM, and low latency of GC, we propose hybrid LLC designs Refresh Energy/bit (pJ) 3.5 - 1.87/ - of GC and STT-RAM. We selected STT-RAM to avoid the high energy overheads imposed by eDRAM. We carried out a design space exploration for the optimal size partitioning between GC and cacheline placement policies or buffering can be used, which can

STT-RAM, while maintaining the same area budget as SRAM, and

potentially increase the number of overlaps. However, we do not found that equal-area (8 MB GC, 1 MB STT-RAM) distribution pursue these optimizations due to the lack of substantial returns on

results in the sweet spot of high performance and low energy. The

either performance or energy. parameters of this organization are listed in Table 6. The ways of 7 HYBRID CACHE HIERARCHY each set of the hybrid cache are split between GC and STT-RAM cachelines, in the ratio of capacity. On a cache lookup, tags of both A growing body of research has proposed either eDRAM or STT- GC and STT-RAM ways are read and compared. If there is a hit in RAM as a replacement for LLCs ([6–8, 16, 18, 39, 45, 53, 69, 75, one of the GC ways, a read or write is carried out. On a miss in GC

77, 88]). In this section, we build on prior work to evaluate hybrid ways, but a hit an STT-RAM way, the cacheline is moved to the LRU cache hierarchies, in an effort to build efficient SRAM “free” on-chip position in the GC ways. Since GC ways tend to have “hot” data, in caches. order to exploit temporal locality, the evicted cacheline from GC way First, we compare proposed GC based caches with other, state-of- is moved to the LRU position of the STT-RAM ways. In case of a

the-art memory technologies. We consider the architectures, where miss in both GC and STT-RAM ways, the cacheline is fetched from L and L caches are kept as GC and use either eDRAM or STT- the next level and placed in the LRU position of STT-RAM ways. RAM in LLC, namely “GC-GC-eDRAM” and “GC-GC-STTRAM” The proposed hybrid cache architecture can be further optimized respectively. STT-RAM parameters were taken from , while via novel replacement policies and prefetchers, which we leave for

parameters for eDRAM are obtained from CACTI simulations, and future work [9, 46]. With L and L cache as GC, and a hybrid LLC, are listed in Table 6. In our experiments, we consider state-of-the-art we evaluate the architecture, results of which are compiled in orange refresh-optimized eDRAM which achieves ~20× reductions in bars of Figures 1 & 16, under “GC-GC-Hybrid”. the number of refreshes over regular eDRAM at 2-3% area overhead. As expected, the energy consumption of hybrid design is very

We compare these technologies with ALL-GC-AREA and present close to that of ALL-GC-AREA: within 2%, on average, across energy and performance comparisons in Figures 1 and 16, re- both single and multi-core benchmarks. More importantly, the in- spectively. These experiments provide several interesting results. creased LLC accesses, due to larger cache, are performed with lower (a) eDRAM based LLC has 2× density benefits over GC, which

As the technology scales down, transistors become leakier with

smaller storage capacitances, resulting in smaller DRTs, exacerbat- ing the refresh problem. This means that at lower technology nodes,

GC based caches, due to a large number of refreshes, could perform

worse than traditional SRAM based caches. To understand scalability characteristics of proposed GCs, we carry out experiments to char- acterize the energy consumption of proposed GC with technology scaling.

We carry out our analysis for 28, 22, 14, 10, and 7 nm tech-

to ALL-SRAM. nology nodes. Generally, leakage current and device capacitance are inversely proportional to technology node [2, 85]. So, scaling latencies of GC. As a result, we observe 5% improvement in IPC down to the next technology node would result in DRT reduction by over ALL-GC-AREA baseline, averaged across multi-core simula- ~50%. Additionally, cell access energies also decrease by ~50% as tions. Compared to the traditional SRAM hierarchy, our proposal we move to lower technology . Considering these trends, we sim-

shows 24% better performance with 42% less energy consumption ulate ALL-SRAM and ALL-GC-AREA configurations and calculate (43% better with 36% less energy, in case of multi-core). These dynamic and refresh energy consumptions. At each technology node, designs can be further optimized by exploiting decoupled bitline we normalize ALL-GC-AREA’s energy consumption (dynamic + optimizations proposed previously, resulting in an extra 13% savings refresh) compared to ALL-SRAM’s energy consumption (dynamic)

in energy, as discussed in Section 6, leading to overall 50% (44% and present the results in Figure 18. Due to space constraints, we in case of multi-core) saving in energy as compared to ALL-SRAM. present only 1 benchmarks, 5 each from different memory inten- In conclusion, while GC works best for L and L caches, real envi- sity groups listed in Table 5. The first five benchmarks are from ronments require large-capacity LLC with low latency, which can MEM_HIGH, followed by MEM_MED and MEM_LOW.

be addressed with our proposed GC-STTRAM hybrid LLC design. We observe that as we move to lower technology nodes, the In the hybrid design, we move the recently accessed cachelines to contribution of refresh energy increases considerably. At 7 nm, the the GC part of hybrid LLC, thus serving them with lower latency, if energy consumption of GC based caches (ALL-GC-AREA) is almost locality exists. comparable to SRAM based caches (ALL-SRAM). Therefore, at

In summary, we present Energy-Delay Product (EDP) results of 7 nm or lower, proposed GC, due to refresh, can perform worse various architectures in Figure 17. As can be observed, for both than SRAM. Many techniques [6, 7, 80, 86] have been explored to single and multi-programmed workloads, any GC based hierarchy reduce refreshes if it becomes a problem. We propose to reduce the does better than the baseline SRAM one. The most favorable design refresh energy by increasing the back-gate bias voltage. Figure 2

point is obtained by utilizing the benefits of asymmetric write - shows reduction in leakage current with increasing back-gate bias optimized GC caches at all levels, and a hydrid STT-RAM - GC voltage in a negative direction. This increases the cell’s DRT, which LLC. This architecture achieves an EDP which is 0.46× of the decreases refresh frequency and hence, energy. In Figure 19, for baseline SRAM one. three technology nodes, we show that the actual DRT of rows can

Figure 18: Energy breakdown (Cache) of ALL-GC-AREA at various technology nodes. Normalized to ALL-SRAM dynamic energy

respectively. However, extending NRP to LLC generates ~12% new misses due to invalidations. It, therefore, results in 23% (36% for multi-programmed workloads) drop in IPC (as seen from Figure 20) as a miss in LLC results in a request to the main memory. Also, while invalidating, we writeback dirty cachelines. This has negli- gible (<10MB/s) bandwidth impact on L and L bandwidth, but

results in an average memory bandwidth usage of ~110MB/s across

multi-core workloads, which is small but may be unacceptable in many cases. Although NRP doesn’t generate many new writebacks

Figure 19: Data Retention Time for different rows, that decides

because it preemptively invalidates cachelines which would other- refresh interval. wise have been dirty evictions. NRP can altogether remove refreshes, and hence refresh energy, from L and L caches, while L cache would still need to be periodically refreshed. However, implement- be much higher than the worst case, for which we have to design ing L with STT-RAM does not need not refresh. As a result, we refresh mechanisms. A potential solution to avoid that could be to conclude that a hybrid cache hierarchy, where the L & L comprise

use architectural solutions like RAIDR and implement separate of proposed GC, and LLC is made from STTRAM-GC hybrid is the refresh intervals for different rows based on their actual DRTs. This lowest energy, highest performance on-chip cache hierarchy with can be achieved by dividing the rows into DRT bins and applying a refresh-free L and L caches. different refresh interval for each bin, leading to a reduction in the number of refresh operations.

9 RELATED WORK

8.1 Refresh Free Hybrid Cache Hierarchy Emerging Memory Technologies for Caches: Due to scalability Another key solution which reduces the effect of refresh is a No and energy issues of traditional SRAM, several studies have been Refresh Policy (NRP), where rather than refreshing a cacheline, we carried out to evaluate emerging memory technologies for caches. invalidate the data if the line has not been touched after a refresh STT-RAM, owing to its low leakage energy and density benefits,

operation has been performed, but before it reaches its DRT. The has been viewed as a promising candidate. However, it suffers from key observation here is that if a row is accessed, the countdown for inherent weaknesses - high write latency and write energy. There its DRT is reset, starting the refresh window from that point. NRP have been many proposals to alleviate these shortcomings [8, 18, 34, can further be optimized by taking advantage of the high number 37, 45, 69, 70, 75, 77, 82, 93].

of writes in caches to reduce such invalidations (similar to ). Another potential replacement for SRAM are eDRAMs, which To invalidate GC cachelines, a dedicated, on-chip refresh controller, offer high density, low leakage, similar access latencies, and low which generates a pulse at given intervals would be required [7, 31]. dynamic energies. However, eDRAM requires refresh operations to We propose to maintain a 5-bit (programmable) saturating counter preserve data integrity . As cache size increases, each refresh

for every cacheline, which is incremented every 132th epoch of the requires more energy, and more lines need to be refreshed; thus, refresh interval. Once the counter saturates, the cacheline is invali- refresh can potentially become the main source of eDRAM power dated. If the cacheline was dirty, it is written back to the next level of dissipation. Many studies have been carried out to amortize this the hierarchy. In case of a write to the cacheline, the counter is reset effect [7, 16, 31, 56, 80, 86, 87]. For instance, showed that the

to 0. We implement this policy at 28nm and calculate the number DRTs of cells in large eDRAM modules exhibit spatial correlations, of times a cacheline was invalidated due to DRT expiration and and exploit this behavior to reduce refresh energy. In contrast, the present the percentage of cache read misses due to such invalidations proposed GC already has insignificant refresh overheads. on the y1-axis of Figure 2 and the performance impact of these Gain Cell: Many circuit and array level architectures have been

invalidations on the y2-axis. We observe that NRP in L and L proposed for GCs [10, 21–23, 28–30, 42, 48, 76]. The biggest ad- caches does not cause many misses (<1%), and exhibits similar vantage of the GC is its logic compatible fabrication process - only performance as compared to GC with refreshes - on average, <0.1% transistors are used for designing the cell and, therefore, can use and ~1% drop in IPC, for single and multi-programmed workloads the same fabrication process as the processor. Transistor’s parasitic

Figure 20: Percentage of Reads that miss due to No Refresh Policy (NRP) invalidations (y1-axis). Performance of NRP, normalized

capacitor is used to store the data. Due to low storing capacitance for removing refreshes in GC caches closer to the CPU. Finally, and high leakage of the junctions in bulk-MOSFET , DRT of we show that proposed GC, in conjunction with emerging memory GC is tiny. To improve DRT, Robert et al. propose using FDSOI- technologies like STT-RAM can be used to architect SRAM-free MOSFET instead of bulk-MOSFET or FinFET, which has reduced cache hierarchies with a much superior energy-delay product as

leakage current as the junctions are isolated in FDSOI through the compared to SRAM caches. oxide . The other added advantage of the FDSOI device is its back-gate bias feature . While body-bias can also be applied in REFERENCES bulk devices to improve the DRT , it has limited breakdown volt- [n. d.]. FDSOI 2 nm Technology. http://www.st.com/b/content/st_com/en/about/ age and can only be applied in p-type MOSFETs . In contrast, a innovation---technology/FD-SOI.html.

[n. d.]. ITRS Scaling. http://www-inst.eecs.berkeley.edu/~ee130/sp06/chp7.pdf. large back-gate bias can be applied in FDSOI technology. Micron Technical Note TN-41-01. http://www.micron.com/products/ In the FDSOI device, threshold voltage and leakage of the device support/power-calc/. can be changed dynamically with the help of the back-gate bias. 2018. 12FDX. https://www.globalfoundries.com/technology-solutions/cmos/fdx/ 12fdx/.

Many designs [61, 81, 89, 90] have been proposed by exploiting the James W Adkisson, Ramachandra Divakaruni, Jeffrey P Gambino, and Jack A back-gate bias feature of the FDSOI. designed the power gating Mandelman. 2002. Embedded DRAM on silicon-on-insulator substrate. US Patent 6,350,653. circuit with the help of the back-gate bias feature of the FDSOI Aditya Agrawal, Amin Ansari, and Josep Torrellas. 2014. Mosaic: Exploiting the device. Clerc et al. designed the robust multiplier and DC-to- spatial locality of process variation to reduce refresh energy in on-chip eDRAM

DC converter by exploiting the back-gate bias feature. In , GC is modules. In Proceedings of the 20th International Symposium on High Perfor- mance Computer Architecture (HPCA). implemented using FDSOI-MOSFET for its low leakage feature and Aditya Agrawal, Prabhat Jain, Amin Ansari, and Josep Torrellas. 2013. Refrint: uses 2 additional transistors to improve DRT at the cost of the area. Intelligent Refresh to Minimize Power in On-chip Multiprocessor Cache Hier-

Architectural optimizations for SRAM caches: Many schemes archies. In Proceedings of 19th International Symposium on High Performance Computer Architecture (HPCA). have been looked into to reduce SRAM’s leakage energy [25, 41]. Junwhan Ahn, Sungjoo Yoo, and Kiyoung Choi. 2014. DASCA: Dead write predic- To catch up with the increasing demands for large capacity caches, tion assisted STT-RAM cache architecture. In Proceedings of 20th International

Symposium on High Performance Computer Architecture (HPCA). numerous cache compression techniques have been looked into [12, Junwhan Ahn, Sungjoo Yoo, and Kiyoung Choi. 2015. Prediction hybrid cache: 60, 64]. Also, there has been a significant amount of work on An energy-efficient STT-RAM cache architecture. IEEE Trans. Comput. (2015). achieving higher performance in caches. Studies have been carried Yoshiyuki Ando. 2003. Capacitorless DRAM gain cell. US Patent 6,560,142.

Jeff Andrews and Nick Baker. 2006. Xbox 3 system architecture. IEEE micro

out for efficient cache replacement and cache management poli- 26, 2 (2006), 25–37. cies [17, 35, 58, 66, 67], dead block predictions [44, 50, 57], or ex- Angelos Arelakis, Fredrik Dahlgren, and Per Stenstrom. 2015. Hycomp: A Hybrid ploiting the differences between reads and writes in caches [43, 83]. Cache Compression Method for Selection of Data-type-specific Compression

Methods. In Proceedings of the 48th International Symposium on Microarchitec-

ture (MICRO). 1 CONCLUSION E. Ashenafi and M. H. Chowdhury. 2018. A New Power Gating Circuit Design

Approach Using Double-Gate FDSOI. IEEE Transactions on Circuits and Systems

In this work, we propose a novel FDSOI based 2T Gain Cell (GC) II: Express Briefs 65, 8 (2018), 1074–1078. as a promising candidate for use in all levels of on-chip caches. Christian Bienia, Sanjeev Kumar, Jaswinder Pal Singh, and Kai Li. 2008. The

PARSEC Benchmark Suite: Characterization and Architectural Implications. In

The proposed GC has better energy efficiency, a much smaller area Proceedings of 17th International conference on Parallel Architectures and Com- footprint, and better scalability as compared to 6T SRAM. Using pilation Techniques (PACT). the back-gate bias feature of FDSOI, we improve GC’s data reten- Trevor E Carlson, Wim Heirman, and Lieven Eeckhout. 2011. Sniper: Exploring the Level of Abstraction for Scalable and Accurate Parallel Multi-core Simulation.

tion time, minimizing contribution of refresh to the overall energy In Proceedings of International Conference for High Performance Computing, consumption of the cache hierarchy. Networking, Storage and Analysis (SC). We evaluate various architectural implementations of GC, for Mu-Tien Chang, Paul Rosenfeld, Shih-Lien Lu, and Bruce Jacob. 2013. Tech- nology comparison for large last-level caches (L3Cs): Low-leakage SRAM, low

all levels of on-chip caches and demonstrate that GC based caches write-energy STT-RAM, and refresh-optimized eDRAM. In Proceedings of 19th substantially reduce the dynamic energy consumption of memory International Symposium on High Performance Computer Architecture (HPCA).

M. Chaudhuri. 2009. Pseudo-LIFO: The Foundation of a New Family of Re-

subsystem as compared to traditional SRAM caches. We exploit placement Policies for Last-level Caches. In Proceedings of the 42nd Annual the inherent capabilities of the proposed GC, including decoupled IEEE/ACM International Symposium on Microarchitecture (MICRO). read and write bitlines, to further reduce dynamic energy. Next, Xunchao Chen, Navid Khoshavi, Jian Zhou, Dan Huang, Ronald F DeMara,

Jun Wang, Wujie Wen, and Yiran Chen. 2016. AOS: Adaptive Overwrite

we demonstrate that a no-refresh policy, where GC based cache Scheme for Energy-Efficient MLC STT-RAM cache. In Proceedings of the 53rd lines are invalidated than refreshed, can be a viable implementation ACM/EDAC/IEEE Design Automation Conference (DAC).

HyGain: High Performance, Energy-Efficient Hybrid Gain Cell based Cache Hierarchy

Kangguo Cheng, Ali Khakifirooz, Kern Rim, and Ramachandra Divakaruni. 2015. SH Kang and C Park. 2017. MRAM: Enabling a sustainable device for pervasive Method and Structure For Forming a Localized SOI finFET. US Patent 8,987,823. system architectures and applications. In Proceedings of the IEEE International A. Chhabra and V. Rana. 2016. -1.1V to +1.1V 3:1 Power Switch Architecture Electron Devices Meeting (IEDM). for Controlling Body Bias of SRAM Array in 28nm UTBB CMOS FDSOI. In E. Karl, Z. Guo, J. W. Conary, J. L. Miller, Y. Ng, S. Nalam, D. Kim, J. Keane,

20 29th International Conference on VLSI Design and 20 15th International U. Bhattacharya, and K. Zhang. 2015. 17.1 A 0.6V 1.5GHz 84Mb SRAM design Conference on Embedded Systems (VLSID). 179–184. in 14nm FinFET CMOS technology. In Proceedings of the IEEE International Woong Choi, Gyuseong Kang, and Jongsun Park. A refresh-less eDRAM Solid-State Circuits Conference - (ISSCC) Digest of Technical Papers. 1–3.

macro with embedded voltage reference and selective read for an area and power Stefanos Kaxiras, Zhigang Hu, and Margaret Martonosi. 2001. Cache decay: efficient Viterbi decoder. IEEE Journal of Solid-State Circuits 50, 1 (2015), exploiting generational behavior to reduce cache leakage power. In Proceedings 2451–2462. 28th International Symposium on Computer Architecture (ISCA).

Ki Chul Chun, Pulkit Jain, Tae-Ho Kim, and Chris H Kim. 2012. A 6 MHz Amit Kazimirsky, Adam Teman, Noa Edri, and Alexander Fish. 2017. A 0.65-v, Logic-Compatible Embedded DRAM Featuring an Asymmetric 2T Gain Cell for 500-mhz integrated dynamic and static ram for error tolerant applications. IEEE High Speed On-die Caches. IEEE Journal of Solid-State Circuits 47, 2 (2012), Transactions on Very Large Scale Integration (VLSI) Systems 25, 9 (2017), 2411–

547–559. 2418. Ki Chul Chun, Pulkit Jain, Jung Hwa Lee, and Chris H Kim. 2011. A 3T Gain S. Khan, A. R. Alameldeen, C. Wilkerson, O. Mutluy, and D. A. Jimenezz. 2014. Cell Embedded DRAM Utilizing Preferential Boosting for High Density and Improving Cache Performance Using Read-Write Partitioning. In Proceedings of Low Power on-die Caches. IEEE Journal of Solid-State Circuits 46, 6 (2011), the 20th International Symposium on High Performance Computer Architecture

1495–1505. (HPCA). S. Clerc, M. Saligane, F. Abouzeid, M. Cochet, J. Daveau, C. Bottoni, D. Bol, J. Samira Manabi Khan, Yingying Tian, and Daniel A Jimenez. 2010. Sampling Dead De-Vos, D. Zamora, B. Coeffic, D. Soussan, D. Croain, M. Naceur, P. Schamberger, Block Prediction for Last-Level Caches. In Proceedings of the 43rd International P. Roche, and D. Sylvester. 2015. 8.4 A 0.33V/-40C Process/Temperature Closed- Symposium on Microarchitecture (MICRO).

Loop Compensation SoC Embedding all-Digital Clock Multiplier and DC-DC Navid Khoshavi, Xunchao Chen, Jun Wang, and Ronald F DeMara. 2016. Read- converter exploiting FDSOI 28nm back-gate biasing. In 20 IEEE International tuned stt-ram and edram cache hierarchies for throughput and energy enhancement. Solid-State Circuits Conference - (ISSCC) Digest of Technical Papers. 1–3. arXiv preprint arXiv:1607.080 (2016). Krisztián Flautner, Nam Sung Kim, Steve Martin, David Blaauw, and Trevor Namhyung Kim, Junwhan Ahn, Kiyoung Choi, Daniel Sanchez, Donghoon Yoo,

Mudge. 2002. Drowsy caches: simple techniques for reducing leakage power. and Soojung Ryu. 2018. Benzene: an energy-efficient distributed hybrid cache In Proceedings 29th Annual International Symposium on Computer Architecture architecture for manycore systems. ACM Transactions on Architecture and Code (ISCA). Optimization (TACO) (2018). Eric J Fluhr, Joshua Friedrich, Daniel Dreps, Victor Zyuban, Gregory Still, Christo- Toshiaki Kirihata, Paul Parries, David R Hanson, Hoki Kim, John Golz, Gregory

pher Gonzalez, Allen Hall, David Hogenmiller, Frank Malgioglio, Ryan Nett, et al. Fredeman, Raj Rajeevakumar, John Griesemer, Norman Robson, Alberto Cestero, 2014. 5.1 POWER TM: A 12-Core Server-Class Processor in 22nm SOI with et al. 2005. An 800-MHz embedded DRAM with a concurrent refresh mode. 7.6 Tb/s off-chip Bandwidth. In Proceedings of the 20 IEEE International IEEE Journal of Solid-State Circuits 40, 6 (2005), 1377–1387. Solid-State Circuits Conference Digest of Technical Papers (ISSCC). Wolfgang H Krautschneider and Werner M Klingenstein. 1994. Process for the

Mrinmoy Ghosh and Hsien-Hsin S Lee. 2007. Smart Refresh: An enhanced Manufacture of a High Density Cell Array of Gain Memory Cells. US Patent memory controller design for reducing energy in conventional and 3D die-stacked 5,308,783. DRAMs. In Proceedings of the 40th Annual IEEE/ACM International Symposium Emre Kültürsay, Mahmut Kandemir, Anand Sivasubramaniam, and Onur Mutlu. on Microarchitecture (MICRO). 2013. Evaluating STT-RAM as an energy-efficient main memory alternative. In

R. Giterman, A. Fish, A. Burg, and A. Teman. 2018. A 4-Transistor nMOS-Only Proceedings of International Symposium on Performance Analysis of Systems and Logic-Compatible Gain-Cell Embedded DRAM With Over 1.6-ms Retention Software (ISPASS). Time at 7 mV in 28-nm FD-SOI. IEEE Transactions on Circuits and Systems I: An-Chow Lai, Cem Fide, and Babak Falsafi. 2001. Dead-block prediction & Regular Papers 65, 4 (2018), 1245–1256. dead-block correlating prefetchers. In Proceedings 28th Annual International

R. Giterman, A. Fish, N. Geuli, E. Mentovich, A. Burg, and A. Teman. 2018. An Symposium on Computer Architecture (ISCA). 800-MHz Mixed-VT 4T IFGC Embedded DRAM in 28-nm CMOS Bulk Process Leslie Lamport. 1986. LATEX: A Document Preparation System. Addison-Wesley, for Approximate Storage Applications. IEEE Journal of Solid-State Circuits 53, 7 Reading, MA. (2018), 2136–2148. Donghyuk Lee, Yoongu Kim, Vivek Seshadri, Jamie Liu, Lavanya Subramanian,

R. Giterman, A. Teman, P. Meinerzhagen, L. Atias, A. Burg, and A. Fish. 2016. and Onur Mutlu. 2013. Tiered-latency DRAM: A low latency and low cost DRAM Single-Supply 3T Gain-Cell for Low-Voltage Low-Power Applications. IEEE architecture. In Proceedings of the 20 IEEE 19th International Symposium on Transactions on Very Large Scale Integration (VLSI) Systems 24, 1 (2016), 358– High Performance Computer Architecture (HPCA). 362. K Lee, R Chao, K Yamane, VB Naik, H Yang, J Kwon, NL Chung, SH Jang, B

Nagendra Gulur, R Govindarajan, and Mahesh Mehendale. 2016. MicroRefresh: Behin-Aein, JH Lim, et al. 2018. 22-nm FD-SOI Embedded MRAM Technology Minimizing refresh overhead in DRAM caches. In Proceedings of the Second for Low-Power Automotive-Grade-l MCU Applications. In Proceedings of the International Symposium on Memory Systems. IEEE International Electron Devices Meeting (IEDM). Per Hammarlund, Alberto J Martinez, Atiq A Bajwa, David L Hill, Erik Hallnor, Junmin Lin, Yu Chen, Wenlong Li, Aamer Jaleel, and Zhizhong Tang. 2009. Under-

Hong Jiang, Martin Dixon, Michael Derr, Mikal Hunsaker, Rajesh Kumar, et al. standing the memory behavior of emerging multi-core workloads. In Proceedings 2014. Haswell: The fourth-generation Intel Core Processor. IEEE Micro 34, 2 of the Eighth International Symposium on Parallel and Distributed Computing, (2014), 6–20. ISPDC’09. Lisa Hsu, Ravi Iyer, Srihari Makineni, Steve Reinhardt, and Donald Newell. 2005. Jamie Liu, Ben Jaiyen, Richard Veras, and Onur Mutlu. 2012. RAIDR: Retention-

Exploring the cache design space for large scale CMPs. ACM SIGARCH Computer aware intelligent DRAM refresh. In ACM SIGARCH Computer Architecture News, Architecture News 33, 4 (2005), 24–33. Vol. 40. IEEE Computer Society, 1–12. Mohsen Imani, Abbas Rahimi, Yeseong Kim, and Tajana Rosing. 2016. A low- Prasanth Mangalagiri, Karthik Sarpatwari, Aditya Yanamandra, VijayKrishnan power hybrid magnetic cache architecture exploiting narrow-width values. In Narayanan, Yuan Xie, Mary Jane Irwin, and Osama Awadel Karim. 2008. A

Proceedings of the 5th Non-Volatile Memory Systems and Applications Symposium low-power phase change memory based hybrid cache architecture. In Proceedings (NVMSA). of the 18th ACM Great Lakes symposium on VLSI. ACM, 395–398. Aamer Jaleel, Kevin B. Theobald, Simon C. Steely, Jr., and Joel Emer. 2010. R Manikantan, Kaushik Rajan, and Ramaswamy Govindarajan. 2011. NUcache:

High Performance Cache Replacement Using Re-reference Interval Prediction An efficient multicore cache organization based on next-use distance. In Proceed- (RRIP). In Proceedings of the 37th Annual International Symposium on Computer ings of the IEEE 17th International Symposium on High Performance Computer Architecture (ISCA). Architecture. IEEE, 243–253. Naifeng Jing, Yao Shen, Yao Lu, Shrikanth Ganapathy, Zhigang Mao, Minyi Raman Manikantan, Kaushik Rajan, and Ramaswamy Govindarajan. 2012. Prob-

FAQ

Cadence Virtuoso, Spectre, Innovus, Assura/Quantus; Siemens EDA; Synopsys; Keysight ADS; LTspice.
Yes — design notes, simulation guidance, report, PPT and viva Q&A.