Hardware Implementation For Fast Block Generator
Of Litecoin Blockchain System
Le Vu Trung Duong1,2 , Doan Van Hieu1,2 , Pham Hoai Luan3 , Tran Thi Hong3 , Lam Duc Khai1,2
1
2021 International Symposium on Electrical and Electronics Engineering (ISEE) | 978-1-6654-1487-6/21/$31.00 ©2021 IEEE | DOI: 10.1109/ISEE51682.2021.9418691
University of Information Technology, Ho Chi Minh City, Vietnam
2
Vietnam National University, Ho Chi Minh City, Vietnam
3
Nara Institute of Science and Technology, Japan
Email: 15520146@gm.uit.edu.vn, doanhieutlvn@gmail.com, pham.hoai luan.ox7@is.naist.jp, hong@is.naist.jp, khaild@uit.edu.vn
Abstract—The development of blockchain-based cryptocurrencies has attracted a lot of interest in many fields, especially
in academic research. Besides Bitcoin, Litecoin is one of the
most popular cryptocurrency based on blockchain currently. The
Litecoin mining process uses the Scrypt algorithm as a hash
function in Proof-of-Work (PoW) consensus mechanism. The
yield of cryptocurrency mining depends a lot on the processing
performance of the hardware. Therefore, this paper proposes
a hardware design of the Scrypt algorithm in the Litecoin
mining system using the Verilog hardware description language
(HDL). In the proposed design, the pipeline technique is applied
at two levels: the overall Scrypt level and the PBKDF2 level.
Additionally, the design is divided into several stages to reduce the
latency of each clock cycle and increase the operating frequency.
The simulation and verification of the design are performed
based on Atlys Spartan-6 FPGA Board and Xilinx Virtex-7
FPGA VC707. The performance of the design evaluated on Atlys
Spartan-6 has the hash rate of 2209 hashes/s. The maximum
operating frequency of the system is 147.059MHz. On Xilinx
Virtex-7, the hash rate of the design reaches 5332 hashes/s and
the maximum operating frequency reaches 355.051 MHz.
Index Terms—Blockchain, Litecoin, Mining, Scrypt.
I. I NTRODUCTION
In recent years, the popularity of digital currencies known
as cryptocurrencies have attracted intense interest in both
research and application in the field of finance, economy,
academia. The Bitcoin electronic cash system is the first cryptocurrency based on blockchain technology. Blockchain can be
considered as a massive ledger of cryptocurrency transactions
built into a distributed and decentralized scheme. Blockchain’s
operations are based on peer-to-peer (P2P) networks, economic incentive mechanisms, cryptographic techniques, and
distributed consensus mechanisms. Nodes in P2P network
joining in transaction validation and block creation are called
miners. Miners continuously process block header with the
hash function until the blockchain’s requirement is met. This
process is called mining [1]. Miners participating in the mining
process need to solve a resource-intensive task based on the
PoW consensus mechanism [2]. The reward for miners who
successfully perform PoW-based mining is cryptocurrencies.
The fast growth of Bitcoin has motivated other cryptocurrencies to be born, including Litecoin - a fork of Bitcoin. Litecoin
is created and increasingly popular with advantages compared
978-0-7381-3196-2/21/$31.00 ©2021 IEEE
to Bitcoin such as lower average block creation time (2.5
minutes) and more maximum amount of coins (84 million) [3].
As an alternative coin of Bitcoin, Litecoin has many inherited
characteristics from Bitcoin, such as block structure, PoW
consensus mechanism. The main difference between them is
the hash function required in PoW. Litecoin uses Scrypt and
Bitcoin uses SHA256. Litecoin mining is based on PoW using
the Scrypt as a hash algorithm. The Scrypt is a memory-hard
algorithm, which means that the processing of Scrypt needs
a speed of memory access rather than speed of computing
[4]. Therefore, it is necessary to build specialized hardware to
handle Scrypt to enhance miners’ mining performance.
In paper [5], the author proposed a Litecoin mining system based on FPGA using pipeline architecture. The use of
pipeline techniques reduces 7.92% of the clock cycles of this
design, but its operating frequency is low. The proposed design
is optimized using a pipeline technique with two memories and
divided into multiple processing stages to increase the overall
operating frequency.
The rest of the paper is organized as follows: Section
II provides the background of blockchain, mining, and the
Scrypt hash algorithm. In section III, the hardware architecture
for the design of the Scrypt hash algorithm is presented in
detail. Simulation and implementation results are mentioned
in section IV. Finally, the conclusion of this paper is given in
section V.
II. BACKGROUND
A. Blockchain and Mining
1) Blockchain: Blockchain is a data structure composed of
many linked blocks in order. The following block is linked to
the previous block by its hash value of the block header [2].
The structure of each block consists of two parts: block header
and block body. The block body contains the hash values of
the verified transactions (capacity up to 1MB) and a value
representing the total number of transactions in that block
(transaction counter). The block header contains the fields used
for block validation. These fields are defined as follows:
•
Version (32-bit): Litecoin blockchain protocol version.
9
Authorized licensed use limited to: Univ of Calif Santa Barbara. Downloaded on June 16,2021 at 17:02:03 UTC from IEEE Xplore. Restrictions apply.
Algorithm 1 DK = Scrypt(message, salt, N, p, r)
Parameters for Litecoin:
message = salt = BlockHeader (640 bits)
Block size factor (r) = 1
Number of iterations (c) = 1
Parallelization parameter (p) = 1
CPU/Memory cost parameter (N) = 1024
Length of DK in bits (dklen) = 256
Steps:
1: X ← PBKDF2(message, salt, c, 1024×r×p)
2: Y ← Romix(X, N)
3: DK ← PBKDF2(message, Y, c, dklen)
4: return DK
Algorithm 2 DK = PBKDF2(message, salt, c, dklen)
1: for i ← 116 to (dklen/256)16 do
2:
DKc ← HMAC({salt, i}, message)
3:
DK ← {DK, DKc }
4: end for
5: return DK
Previous block hash (256-bit): hash value of previous
block.
• Merkle root (256-bit): the hash value represents all the
transactions included in the block.
• Timestamp (32-bit): the time when block is created (start
from 1970-01-01T00:00 UTC).
• Bits (32-bit): target difficulty for finding Nonce in PoW.
32
• Nonce (32-bit): a variable in the range of 0 to 2 -1.
2) Mining: Litecoin mining is based on PoW that requires
participants to solve a complex problem using hardware power
[2]. Fig. 1 illustrates a basic model of the Litecoin mining
system. This system includes one Scrypt hardware module
taking BlockHeader as input. The output of Scrypt (B) is
compared with Bits value (A). If B is higher than A, the hash
output (HashOut) meets the Litecoin system’s difficulty, and
the current nonce is determined to be the Golden Nonce. In
contrast, the nonce is increased by one by the Nonce Generator
and processed by Scrypt again until the comparison condition
is satisfied.
•
BlockHeader
Version
Previous
hash
Merkle
root
Time
stamp
nBits
Nonce
Scrypt
Hardware
HashOut
A
B
Comparator
Controller
enable
store
enable
Nonce
Generator
Nonce
Register
Golden
Nonce
Figure 1. Basic Model of Litecoin Mining System
Algorithm 3 hash = HMAC(message, salt)
1: IPAD ← 512 bits of 36363636...3616
2: OPAD ← 512 bits of 5C5C5C5C...5C16
3: KHASH ← SHA256(salt)
4: IXor ← KHASH ⊕ IPAD
5: OXor ← KHASH ⊕ OPAD
6: IHASH ← SHA256({IXor, message})
7: OHASH ← SHA256({OXor, IHASH})
8: hash ← OHASH
9: return hash
Algorithm 4 digest = SHA256(message in)
1: (M, N) ← Padding(message in)
2: H(0) ← H Constants
3: for t ← 0 to (N-1) do
4:
W ← BlockDecomposition(M(t) )
5:
H(t+1) ← HashComputation(H(t) , K Constants, W)
6: end for
7: return digest ←
(N ) (N ) (N ) (N ) (N ) (N ) (N ) (N )
{H1 ,H2 ,H3 ,H4 ,H5 ,H6 ,H7 ,H8 }
B. Scrypt hash algorithm
The Scrypt algorithm is introduced as a password-based key
derivation function by Colin Percival in the paper [4]. It is a
sequential memory-hard function that is initially constructed
against attacks using custom hardware. Parameters required in
Scrypt can be adjusted according to the user’s intent depending
on the amount of memory, available computing power, and
other factors. The creator of Litecoin uses a simple parameter
set of (r, N, p) corresponding to (1, 1024, 1) [3]. With
these parameters, 128KB memory for a single computation
is required. Scrypt algorithm has three execution phases,
including two phases in PBKDF2 and one phase in Romix
[6]. The Alg. 1 describes the Scrypt algorithm customized to
the Litecoin system’s requirement.
1) PBKDF2: The Password-Based Key Derivation Function 2 (PBKDF2) uses the Hash-based Message Authentication
Code (HMAC) to generate a derived key [7]. The size of the
generated key depends on the parameters from Scrypt. The
Alg. 2 shows the pseudocode of PBKDF2 based on Scrypt’s
parameter (parameter c is set to 1 as hard-coded value).
HMAC is a message authentication function with a cryptographic hash function and a secret cryptographic key [8]. It is
used to verify the data integrity or authenticity of messages.
Detailed HMAC is presented in Alg. 3.
Secure Hash Algorithm 2 (SHA2) is a set of cryptographic
hash functions designed by the United States National Security Agency [9]. In PBKDF2, SHA256, which is a SHA2
family, is used in HMAC function as a cryptographic hash
function. As shown in Alg. 4, SHA256 consists of 3 steps:
Padding, BlockDecomposition, HashComputation. The initial
hash values called H constants is a sequence of 32 bits, which
are the fractional parts of the square roots of the first eight
prime numbers. The K constants is a sequence of 32 bits,
10
Authorized licensed use limited to: Univ of Calif Santa Barbara. Downloaded on June 16,2021 at 17:02:03 UTC from IEEE Xplore. Restrictions apply.
Algorithm 5 W = BlockDecomposition(block)
1: for i ← 0 to 63 do
2:
if i<16 then
3:
Wi = block[(32 × i):(32 × i) + 31]
4:
else
5:
Wi = σ1 (Wi−2 ) + Wi−7 + σ0 (Wi−15 ) + Wi−16
6:
end if
7: end for
8: return W
Algorithm 6 H = Hash Computation(H, K, W)
1: (a, b, c, d, e, f, g, h)←(H1 , H2 , H3 , H4 , H5 , H6 , H7 , H8 )
2: for i ← 0 to 63 do
3:
T1 ← h + Σ1 (e) + Ch(e,f,g) + Ki + Wi
4:
T2 ← Σ0 (a) + Maj(a,b,c)
5:
(a, b, c, d, e, f, g, h) ← (T1 +T2 , a, b, c, d+T1 , e, f, g)
6: end for
7: (H1 , H2 , H3 , H4 , H5 , H6 , H7 , H8 ) ←
(H1 +a, H2 +b, H3 +c, H4 +d, H5 +e, H6 +f, H7 +g, H8 +h)
8: return H
which are the fractional parts of the cube roots of the first 64
prime numbers. At Padding step, message in is decomposed
into many 512-bit blocks. The last block is appended with a
padded string. The last block’s length is less than 447 bits
because the padding string’s minimum length is 65 bits. The
Alg. 5 describes Block Decomposition step in detail, where N
is the number of blocks after padding input message. Some
special functions in Alg. 5 and Alg. 6 are described below.
The function RotR(X, k) rotates X to the right by k bits.
σ0 (X) = RotR(X, 7) ⊕ RotR(X, 18) ⊕ (X >> 3)
σ1 (X) = RotR(X, 17) ⊕ RotR(X, 19) ⊕ (X >> 10)
Ch(X, Y, Z) = (X|Y ) ⊕ (X|Z)
M aj(X, Y, Z) = (X|Y ) ⊕ (X|Z) ⊕ (Y |Z)
Σ0 (X) = RotR(X, 2) ⊕ RotR(X, 13) ⊕ RotR(X, 22)
Σ1 (X) = RotR(X, 6) ⊕ RotR(X, 11) ⊕ RotR(X, 25)
2) ROMIX: The Romix is a sequential memory-hard function that defined in Alg. 7. This algorithm fills the memory
of size N×128×r bytes with pseudorandom values handled by
Blockmix in ascending order. Then, it accesses those values in
a pseudorandom order (depending on the result of integerify
function). Therefore, it requires a lot of memory and makes
its implementation in parallel inefficient [6].
BlockMix algorithm is defined in Alg. 8 as a subroutine
of Romix. The Blockmix mixes data from the previous phase
using Salsa20/8 core. Salsa20 is a hash function that takes
a 16×32 bits string as input and produces 16×32 bits. The
Salsa20/8 is a variant of Salsa20 family stream cipher reduced
from 20 rounds to 8 rounds.
The Salsa20/8 core can be defined mathematically as follows, where X = (x0 , x1 , x2 , ..., x15 ) as 16 words of 32 bits
in little-endian format [10]:
Salsa20/8(X) = X+DoubleRound8/2 (X)
DoubleRound(X) = RowRound(ColumnRound(X)).
That means that there are a total of 8 rounds, including
Algorithm 7 B’ = Romix(block)
1: for i ← 0 to (N−1) do
2:
Memi ← block
3:
block ← Blockmix(block)
4: end for
5: for i ← 0 to (N−1) do
6:
j ← Integerify(block) mod N
(where Integerify(Q[0],..,Q[2×r−1]) is defined as the
integer Q[2×r−1] in little-endian format)
7:
block ← Blockmix(block⊕Memj )
8: end for
9: block’ ← block
10: return block’
Algorithm 8 block’ = Blockmix(block)
1: X ← block1 // block = {block1 , block0 }
2: for i ← 0 to (2×r−1) do
3:
X ← X ⊕ block0
4:
X ← X + Salsa20/8(X)
5:
block’i ← X
6: end for
7: return {block’0 , block’1 }
four rounds of RowRound (in short, RR) and four rounds
of ColumnRound (in short, CR). The detailed Salsa20/8 core
description is shown below, where RotL is left rotation by
k-bit operator [11].
• Define QuarterRound (in short, QR):
If T = (t0 , t1 , t2 , t3 )
then QR(T ) = (w0 , w1 , w2 , w3 ) where:
w1 = t1 ⊕ RotL[(t0 + t3 ), 7];
w2 = t2 ⊕ RotL[(t1 + t0 ), 9];
w3 = t3 ⊕ RotL[(t2 + t1 ), 13];
w0 = t0 ⊕ RotL[(t3 + t2 ), 18];
• If X = (x0 , x1 , x2 , ..., x15 )
then CR(X) = (y0 , y1 , y2 , ..., y15 ) where:
(y0 , y4 , y8 , y12 ) = QR(x0 , x4 , x8 , x12 );
(y5 , y9 , y13 , y1 ) = QR(x5 , x9 , x13 , x1 );
(y10 , y14 , y2 , y6 ) = QR(x10 , x14 , x2 , x6 );
(y15 , y3 , y7 , y11 ) = QR(x15 , x3 , x7 , x11 ).
• If Y = (y0 , y1 , y2 , ..., y15 )
then RR(Y ) = (z0 , z1 , z2 , ..., z15 ) where:
(z0 , z1 , z2 , z3 ) = QR(y0 , y1 , y2 , y3 );
(z5 , z6 , z7 , z4 ) = QR(y5 , y6 , y7 , y4 );
(z10 , z11 , z8 , z9 ) = QR(y10 , y11 , y8 , y9 );
(z15 , z12 , z13 , z14 ) = QR(y15 , y12 , y1 , y14 ).
III. P ROPOSAL OF S CRYPT ALGORITHM HARDWARE
ARCHITECTURE FOR L ITECOIN MINING
A. Scrypt in pipeline
Proposed pipelined Scrypt design is shown in Figure 3. The
areas marked as PBKDF2-1st (for P1 process) and PBKDF22nd (for P2 process) correspond to steps 1 and 3 in Alg. 1.
The remaining area corresponds to step 2.
11
Authorized licensed use limited to: Univ of Calif Santa Barbara. Downloaded on June 16,2021 at 17:02:03 UTC from IEEE Xplore. Restrictions apply.
number of cycles
66563
136339
2910
P1
1st input
2nd input
3rd input
66563
RM_W
66563
RM_R
RM_W
P1
1293
P2
RM_R
RM_W
P1
...
P2
RM_R
P2
Figure 2. Pipeline Diagram of Scrypt
BlockHeader
IHASH
OHASH
IHashOut
RomixIn
1
BlockIn
1
0
BlockDecomposition
BRAM
1
DigestIn 0
HashComputation
MemOut
BLOCMIX
MemOut
0
BRAM
BmixOut
BmixIn
BLOCMIX
1
RomixOut
OXor
IXor
ROMIX-1st
BmixOut
KHASH
BmixIn
PBKDF2-1st
PBKDF2-2nd
Figure 4. Proposed SHA256 Datapath
0
ROMIX-2nd
LUTRAM
BlockHeader
BS
0
“8000...00028016” CC
1
HashOut
Figure 3. Pipelined Scrypt Datapath
CC Concatenation
As mentioned in the previous section, the writing data is
arranged in ascending order of addresses. The reading data
is arranged in a random order of addresses, which makes it
impossible to apply the pipeline with two stages of reading
memory and writing memory to a single core of Romix. Thus,
the Romix module is duplicated and marked as Romix-1st and
Romix-2nd . This is a trade-off between processing speed and
area. Therefore, the proposed Scrypt module is divided into
four pipeline stages as shown in Fig. 3 and Fig. 2: P1, RM W
(writing data to memory), RM R (reading data from memory),
and P2. The Romix-1st and Romix-2nd modules take turns
handling two jobs: RM W and RM R. This means that while
Romix-1st is used for RM R, Romix-2nd is used for RM W.
The pipelined Scrypt module takes 136,339 clock cycles for
the first input. For the 2nd input onwards, the total number of
clock cycles required to generate a hash is 66,563.
B. PBKDF2
According to Alg. 1 and Alg. 2, the PBKDF2 function is
called twice with 2 different values of ”dklenP” parameter.
This only affects the number of times that the HMAC function
is executed. HMAC is decomposed into three subunits that are
KHASH, IHASH, and OHASH corresponding to step 2, 3, 4,
and 5 in Alg. 3. HMAC uses 3 SHA256 modules with the
architecture shown in Fig. 4. Each SHA256 module has three
inputs, including 512-bit BlockIn, 256-bit DigestIn, and 256bit DigestOut. The SHA256 module consists of the BlockDecomposition module (corresponding to step 4 in Alg. 5)
and the HashComputation module (corresponding to step 5 in
adder
DigestOut
BS Bit Split
SHA256
CC
IXor
IPAD = 256 bits
of 363636...3616
OPAD = 256 bits
of 5C5C5C...5C16
CC
OXor
Figure 5. Proposed KHASH Datapath
Alg. 5) executed in parallel. They are performed 64 times, corresponding to 64 iterations. The BlockDecomposition module
and the HashComputation module are devided into 3 stages for
each loop as the second-level pipeline architecture in [12] to
reduce the critical path. Therefore, the SHA256 module takes
a total of 195 clock cycles to process each input completely.
BlockHeader is processed by the KHASH module in both
P1 and P2 processes. The proposed design of KHASH, as
shown in Fig. 5. As SHA256 only processes 512 bits at each
execution time, 640-bit BlockHeader is split into two parts.
The first 512 bits are processed by SHA256 and then stored in
LUTRAM. The last 128 bits containing nonce is concatenated
with padding as an input of SHA256. LUTRAM’s value is
used as the previous hash of SHA256. By using LUTRAM,
195 clock cycles are reduced with each nonce change. The
next step, 256 bits output of SHA256 is bitwise XORed with
IPAD and OPAD. These results are called IXor and OXor as
an output of KHASH.
As shown in Fig. 6, two 4-to-1 selectors marked as P1
and P2 correspond to the call of the PBKDF2 function in
Alg. 1. Four constants ”1”, ”2”, ”3”, ”4” corresponding to
the value of i in the Alg. 2 is selected by the other 4-to-1
selector. As shown in Fig. 7, IHashOut is concatenated with
12
Authorized licensed use limited to: Univ of Calif Santa Barbara. Downloaded on June 16,2021 at 17:02:03 UTC from IEEE Xplore. Restrictions apply.
number of cycles
BlockHeader
BS
10
01
CC
00
ENDIAN
CONV
RomixOut
11
“2”
01
“1”
00
11
CC
01
1
0
SHA256
LUTRAM
1
Figure 6. Proposed IHASH Datapath
0
IHashOut
OXor
CC
1
ENDIAN RomixIn
CV
SHA256
OHASH
IHASH
OHASH
BmixIn MemOut
Salsa20/8 Core
DR
BS
BUFFER
OHASH
IHASH
Figure 8. Pipeline Diagram of 4 Constants in PBKDF2-1st
“8000...00062016”
LUTRAM
975
195
OHASH
IHASH
IHashOut
00
“8000...00030016”
390
1
“3”
“1”
“2”
“3”
“4”
0
10
10
PBKDF2-2nd
“4”
BS
PBKDF2-1st
11
“8000...0004A016”
IXor
195
IHASH
HashOut
0
QR
QR
1
QR
QR
0
QR
QR
QR
QR
BUFFER
RR
CR
(2)
(1)
BmixOut
Figure 7. Proposed OHASH Datapath
”8000...00030016 ” as input for SHA256. Four 256 bits output
of SHA256 in the P1 process are stored in 1024 bits buffer.
The endian converter converts big-endian to little-endian and
vice versa. In the P1 process, IXor and OXor are processed
at the same time and stored in the corresponding LUTRAM.
Then 512-bit BlockHeader is hashed and its result is stored
in LUTRAM. The next 128-bit BlockHeader is concatenated
with the constant ”1” and selected as BlockIn. LUTRAM’s
value is used as DigestIn. The result of SHA256 is transferred
to OHASH and then stored in the first 256-bit of 1024-bit
buffer. This process is repeated until all four constants are
entirely processed. The pipelining technique is applied at this
step to reduce clock cycles. Details of pipeline procedure are
shown in Fig. 8. In the P2 process, IHASH’s processing order
is IXor, the first 512-bit RomixOut, the last 512-bit RomixOut,
and constant ”1” concatenated with padding. It is similar to
P1, but only constant ”1” is processed and passed to OHASH
for the final hash.
For Scrypt pipeline process as shown in Fig. 2, when P2
processes the first input, P1 had processed the 3th input.
Therefore, the IXor and OXor results of the 1st , 2nd and 3rd
input must be stored into LUTRAM with different addresses.
C. Romix
The Romix datapath has two phases of memory (BRAM)
access: writing data to memory (step 1-4 in Alg. 7) and reading
data from memory (step 5-8 in Alg. 7). Each phase consists of
1024 iterations of the Blockmix module. In this design, 128-
Figure 9. Proposed Blockmix Datapath
KB BRAM IP with a single port available in the library of the
Xilinx ISE tool is used.
As shown in Fig. 9, the Blockmix module takes 1024-bit
data as the input and transforms them into 1024-bit output
through the use of 2 iterations of Salsa20/8 core (Alg. 8).
The input data is processed by the xor operator before being
transmitted to the Salsa20/8 core. The buffer is used to store
two 512-bit output values of two Salsa20/8 core iterations
corresponding to 1024 bits of the Blockmix module’s output.
The Salsa20/8 core consisting of 4 iterations of DoubleRound (in short, DR) module coverts 16×32 bits input into
16×32 bits output. Instead of using an unroll style, two
feedback data streams are used (path (1) for Blockmix, path
(2) for Salsa20/8 core, Fig. 9) to saves significantly hardware
resources. In the DR module, there are 8 QuarterRound
(QR) modules divided into two equal parts in parallel: 4 for
ColumnRound (CR) and 4 for RowRound (RR). Every part of
that 4 QR takes three clock cycles because the QR is divided
into three stages, as shown in Fig. 10. This is to reduce the
latency of each cycle and ensure consistency with the SHA256
module.
Specifically, each DR module that consists of 2 arrays of
4 QR modules in serial takes six clock cycles. The Blockmix
module with 2 loops of Salsa20/8 core (equivalent to 2×4 =
8 loops of DR) takes a total of 8×6+16 = 64 clock cycles
(16 clock cycles for intermediate states). As such, Romix
modules take a total of 132,100 clock cycles for 2048 loops
of Blockmix modules at both memory access phases and
intermediate states.
13
Authorized licensed use limited to: Univ of Calif Santa Barbara. Downloaded on June 16,2021 at 17:02:03 UTC from IEEE Xplore. Restrictions apply.
x0
y0
<<<7
<<<9
x1
y1
<<<13
x2
y2
<<<18
x3
y3
stage 1
stage 3
Figure 10. Proposed QuarterRound Datapath
Litecoin blockchain
Block header || Golden nonce
Block hash
testcase_in.txt
expected_out.txt
false
Hardware design in Verilog
Is same?
true
Function is correct
design_out.txt
Figure 11. Verification Flow
IV. D ESIGN S IMULATION AND I MPLEMENTATION
A. Design Simulation
The verification flow is shown in Fig. 11. In particular, block
header and golden nonce values are collected from Litecoin
blockchain [3] and stored in testcase in.txt file as test cases.
Similarly, block hash values are stored in the expected out.txt
file. If the hash output values stored in the design out.txt file
of Scrypt design matches the block hash in expected out.txt,
the function of the proposed design is correct.
B. Design Implementation
After the design’s function is verified, the design synthesis
and implementation are performed by Xilinx ISE Design Suite
14.7 based on Atlys Spartan-6 FPGA Board and Xilinx Virtex7 FPGA VC707 Evaluation Kit.
Table. I shows the comparison between our proposed designs and proposed design in [5]. The design’s performance is
evaluated based on the number of hash calculations performed
in a second called hash rate (hashes per second, H/s).
By dividing into many stages and applying pipeline techniques, the hash rate of the proposed pipelined design reaches
2209 H/s and is 3.6 times higher than [5] and two times
higher than the proposed non-pipelined design. Although the
throughput of [5] is higher than the proposed non-pipelined
design, it is lower than the proposed pipelined design. From
the point of view of the area, the proposed design with pipeline
architecture uses 3.3 times as many slices as [5] and 1.9 times
as many slices as proposed non-pipelined design. Finally,
the design’s implementation result on Virtex-7, the maximum
frequency reaches 355.051 MHz, this speed is higher than
that of Atlys Spartan-6 about 2.141 times. Also, the hash rate
reaches 5332 H/s whereas the area is the same.
Table I
C OMPARISON OF S CRYPT FPGA I MPLEMENTATION R ESULTS
Designs
Design
in [5]
FPGA
Clock Cycles/Hash
Resources(Slices)
Frequency(MHz)
Throughput(Mbps)
Hashrate(H/s)
Spartan-6
17,700
15,983
26.968
1.054
610
Proposed Design
Non-Pipelined
Pipelined
Spartan-6
136,339
27,311
147.059
0.69
1079
Pipelined
Spartan-6
66,563
52,298
147.059
1.414
2209
Virtex-7
66,563
52,298
355.051
3.414
5332
V. C ONCLUSION
In this paper, we have proposed a hardware design of the
Scrypt algorithm that has been optimized for the Litecoin
mining system. The Scrypt hardware design uses a pipeline
technique at a general level with two memories and divided
into multiple processing stages to increase maximum operating
frequency. On the Atlys Spartan-6 FPGA Board, this design
supports the maximum frequency of 147.059MHz and has
a hashing speed of 2209 H/s. Also, the design reaches the
maximum frequency of 355.051 MHz with the hash rate
of 5332 H/s on FPGA Xilinx Virtex-7. With the mentioned
solution, the results show a significant improvement in the
hash rate and the maximum frequency compared with the
conventional works.
ACKNOWLEDGMENT
This research is funded by Vietnam National University
HoChiMinh City (VNU-HCM) under grant number DSC202126-01.
R EFERENCES
[1] Y. Yuan and F. Wang, “Blockchain and cryptocurrencies: Model, techniques, and applications,” IEEE Transactions on Systems, Man, and
Cybernetics: Systems, vol. 48, no. 9, pp. 1421–1428, Sept. 2018.
[2] O. H. J. O. L. Y. Ujan Mukhopadhyay, Anthony Skjellum and R. Brooks,
“A brief survey of cryptocurrency systems,” in 14th Annual Conference
on Privacy, Security and Trust (PST). Auckland, New Zealand: IEEE,
2016, pp. 1–8.
[3] Litecoin wiki. [Online]. Available: https://litecoin.info/
[4] C. Percival, “Stronger key derivation via sequential memory-hard functions,” BSDCan2009, pp. 1–16, May. 2009.
[5] M. Abdillah, I. Hidayat, and R. Ardianto, “Design and implementation
scrypt processor for cryptocurrency with pipeline architecture based on
fpga,” e-Proceeding of Engineering, vol. 2, no. 2, pp. 2075–2082, Aug.
2015.
[6] C. Percival. (2016) The scrypt password-based key derivation function.
RFC 7914.
[7] B. Kaliski, “Pkcs #5: Password-based cryptography specification
version 2.0 (rfc 2898),” IETF, 2000. [Online]. Available:
https://tools.ietf.org/html/rfc2898
[8] The Keyed - Hash Message Authentication Code (HMAC), FIPS PUB
198-1, National Institute of Standards and Technology (NIST) Std., July
2008.
[9] Secure Hash Standard (SHS), FIPS PUB 180-4, National Institute of
Standards and Technology (NIST) Std., August 2015.
[10] D. J. Bernstein. (April 2005) Salsa20 specification. The University of
Illionis, Chicago. [Online]. Available: https://cr.yp.to/suffle/spec.pdf
[11] ——. (April 2005) The salsa20 core. [Online]. Available:
http://cr.yp.to/salsa20.html
[12] L. V. T. Duong, N. T. T. Thuy, and L. D. Khai, “A fast approach for
bitcoin blockchain cryptocurrency mining system,” Integration, vol. 74,
05 2020.
14
Authorized licensed use limited to: Univ of Calif Santa Barbara. Downloaded on June 16,2021 at 17:02:03 UTC from IEEE Xplore. Restrictions apply.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )