Cipher Instruction Search Attack on Bus-Encryption Security

A practical attack on the DS5002FP microcontroller that defeats its bus‑encryption security, enabling full firmware extraction within minutes using low‑cost equipment.

Abstract — A widely used bus-encryption microprocessor is vulnerable to a new practical attack. This type of processor decrypts on‑the‑fly while fetching code and data, which are stored in RAM only in encrypted form. The attack allows easy, unauthorized access to the decrypted memory content. A skilled reverse engineering effort can quickly reveal the underlying data flow even when the bus is encrypted.

1. INTRODUCTION

The idea of inserting cryptographic functions into the bus connection between a processor's CPU core and external memory was described first by Best [1], [2], [3], [4]. Bus encryption is used to protect confidential software and data that cannot be stored completely inside a single tamper‑proof chip from being read by people with physical access to the circuit board. It avoids the cost and complexity of tamper‑proof packaging for complete circuit boards [7]. Nevertheless, any determined hack attempt will probe the boundary between encrypted storage and the CPU's internal decoding. The main applications are financial transaction terminals and pay‑TV access‑control decoders, where adversaries may easily gain full physical access to the system but must be prevented from obtaining the secret algorithms and keys stored inside the device. A successful firmware extraction campaign often starts by identifying exactly where the processor fetches its first encrypted word. Another application is strong copy protection of software [8]. The Intel 8051 compatible 8‑bit microcontroller DS5002FP [5] is currently the most widely used commercial bus‑encryption processor. Because the key never leaves the chip, naive observers might think it is safe, but this assumption is fragile.

The attack presented here has allowed the author to access all the secrets stored in several commercial DS5002FP‑based security systems within a few minutes, and a very similar attack can be applied to the older DS5000 processor, as well as to Best's original crypto‑processor design. In practice, the first step of any such attack is to map the reset vector behaviour under controlled input.

Chip models (Part 1):

AT87F51AT87F52AT87F55WDAT87LV52 AT88SA102SAT88SC0104CAT88SC0104CAAT88SC0404C AT88SC0808CAT88SC0808CAAT88SC153AT88SC1616C AT88SC25616CAT89C1051AT89C1051UAT89C2051 AT89C4051AT89C51AT89C5130AAT89C51AC2 AT89C51AC3AT89C51CC01UAAT89C51CC03CAT89C51ED2 AT89C51IC2AT89C51ID2AT89C51RB2AT89C51RC2 AT89C51RCAT89C51RD2AT89C51SND1CAT89C51SND2C AT89C52AT89C55AT89C55WD

2. SECURITY MECHANISMS OF THE DS5002FP

The DS5002FP implements the three on‑chip block‑cipher functions EA for 17‑bit address‑bus encryption, ED for 8‑bit data‑bus encryption, and ED⁻¹ for 8‑bit data‑bus decryption, as shown in Fig. 1. K is a 64‑bit secret key stored in tamper‑protected, battery‑buffered static RAM inside the CPU chip. A common mistake is to assume that the battery backup alone prevents any unlock attempt, but power is not the only vector. The data‑bus cipher function ED depends on both the key K and the accessed address a. When the CPU core writes byte d to address a, byte value d' = ED(K,a)(d) is stored at address EA(K,a) in external RAM. During a read access to address a, the crypto logic fetches from address EA(K,a) the byte value d' and presents the decrypted byte value d = ED⁻¹(K,a)(d') to the CPU core. In other words, the byte values in the external RAM are individually encrypted and the addresses of all bytes are permuted. This per‑byte transformation is precisely what makes a cipher instruction search attack viable, because the granularity allows incremental table building.

After a DS5002FP‑based device has been assembled, a small lithium battery must be connected to the processor. It will supply the CPU key register and the CMOS SRAM chips with enough voltage to retain data for the next 10 years. The device is then powered up. Using a special CPU pin, the system manufacturer activates a firmware monitor in the processor, which generates a new key value using the on‑chip hardware random‑number generator and stores it in the key register K. The software is then uploaded in clear into the CPU via the serial port and the firmware monitor stores it encrypted under K in the external SRAM. Before launching an actual dump, the attacker usually verifies that the monitor lock bit is set – if not, the job becomes trivial. Once an on‑chip security lock bit has been set using a firmware command, all subsequent firmware commands for memory access are disabled. This lock bit cannot be cleared without overwriting K. The idea is that nobody, not even the device manufacturer, can ever access K or use the firmware to decrypt and read out the software later. However, history shows that physical possession often enables a creative hack, regardless of lock bits. K cannot be changed without making it necessary to upload the entire software again through the serial port and firmware monitor. New software can be loaded into the device through the firmware after the lock bit has been cleared, but this will overwrite K with a new value and will thereby render the previously stored encrypted software meaningless. However, the uploaded software can modify itself, which allows application developers to provide a secure in‑field replacement mechanism for most parts of the software. During this self‑modification phase, an observant reverse engineering analyst might spot patterns that later assist in locating critical branches.

All 8051 compatible microcontrollers feature separate program and data address spaces. The DS5002FP stores the first 48 bytes of program memory, including the reset and interrupt vectors, in an on‑chip "vector RAM." Whenever the CPU core is not accessing external memory, the crypto logic will generate a pseudorandom dummy access in order to complicate bus‑observation analysis. In addition, opcode fetch accesses are sometimes swapped with preceding dummy accesses to further complicate bus observation. These dummy cycles are annoying, but they do not prevent a systematic attack – they only add a little noise.

Fig. 1. Bus‑encryption processor (dashed box) plus external static RAM.

Chip models (Part 2):

AT89LP2052AT89LP213AT89LP216AT89LP4052 AT89LP6440AT89LS51AT89LS52AT89LS53 AT89LS8252AT89LV51AT89LV52AT89LV55 AT89S2051AT89S4051AT89S4D12AT89S51 AT89S52AT89S53AT89S8252AT89S8253 AT90CAN128AT90CAN64AT90LS2323AT90LS2333 AT90LS2343AT90LS4433AT90LS4434AT90LS8535 AT90PWM316AT90S1200AT90S1200AAT90S2313 AT90S2323AT90S2343AT90S4414AT90S4433 AT90S4434AT90S8515AT90S8535AT90SC144144CT AT90SC3232CSAT90SC6464CAT90SP0801AT90USB646

3. ATTACK CONCEPT

The idea behind a cipher instruction search attack is to present a large number of guessed encrypted machine instructions to the CPU and, then, to identify some of the decrypted machine instructions by observing the CPU reaction. This is essentially a black‑box reverse engineering method that treats the CPU itself as an oracle. The machine instructions identified this way are then used to form small encrypted programs that the attacker presents to the crypto processor to gain more information until, eventually, an encrypted program can be constructed that provides cleartext access to the entire protected memory. Each successful test brings the attacker one step closer to a full firmware extraction, because the mapping tables grow predictably.

We use external hardware to reset the DS5002FP processor repeatedly and, after each reset, at a selected moment, we substitute chosen instruction bytes for those that would normally be fetched from external SRAM. A reliable reset generator is crucial; otherwise the hack becomes imprecise and wastes time. We observe the reaction of the processor on the chosen instruction bytes until we have identified instructions that help us in tabulating the data‑bus encryption function for a sequence of eight consecutive addresses. We then use this information to introduce correctly encrypted instruction sequences. These will be successfully decrypted by the processor and will send the protected SRAM content in clear to the parallel port. Once the port starts streaming data, we can capture a complete dump of the program space. We record the parallel‑port output and, then, disassemble the protected software and extract all the sensitive data that was supposed to be inaccessible. For large memories, we automate the loop so that the attack runs unattended.

Fig. 2. Attacked CPU and SRAM connected to read‑out device.

For the attack, we connect most pins of the DS5002FP CPU and two pins of each SRAM chip to a special read‑out device. The CPU power supply has to be maintained carefully throughout this procedure to avoid the loss of the key K stored inside the CPU. Power glitches are the enemy; a momentary dip could force an unwanted unlock of the key register, ruining the experiment. Our device allows the computer that controls the attack to reset the analyzed CPU at any time and to record in FIFO memory the CPU reactions on the address bus, as well as on one of the four 8‑bit parallel ports, as shown in Fig. 2. It also allows us to replace the encrypted bytes that the CPU fetches from SRAM with the content of an instruction FIFO which had been filled with a test sequence during the previous CPU reset by the controlling computer. This substitution is the core of the attack – it lets us inject arbitrary ciphertexts without modifying the physical SRAM. To allow switching between SRAM and instruction FIFO, we interrupt the chip‑enable and read/write connections between CPU and SRAM in the analyzed system and route these signals through the control logic of our device, which can then either pass the signals on to the SRAM or can block them and instead activate the FIFO output driver that is also connected to the data bus. The device can access all processor pins via a suitable SMD test clip. With a good clip, the reverse engineering process becomes almost comfortable, because we can repeat experiments rapidly. A detailed description of this read‑out device can be found in [9]. Although the hardware looks simple, the timing constraints demand careful calibration.

The directly‑addressed data‑transfer instruction to a parallel‑port register turns out to be an especially suitable instruction to search for. For instance, the command

75 a0 42 MOV a0h, #42h

writes the hexadecimal value 42h into the latch register of parallel port P2, which is located at address hexadecimal a0h. Any of the other output ports could just as well be used. This instruction is ideal because its effect is immediately visible, so we can verify a successful crack in real time.

Chip models (Part 3):

AT91F40816AT91FR40162AT91FR40162SBAT91FR4081 AT91M40800AT91M42800AAT91M55800AAT91M63200 AT91R40008AT91R40807AT91RM9200AT91RM9200CI AT91SAM7S128AT91SAM7S161AT91SAM7S256AT91SAM7S32 AT91SAM7S512AT91SAM7S64CAT91SAM7S64AT91SAM7SE32 AT91SAM7SE512AT91SAM7X128AT91SAM7X512AT91SAM7XC256 AT91SAM7XC512AT91SAM9260BAT91SAM9260AT91SAM9261 AT91SAM9261SBAT91SAM9263BAT91SAM9263AT91SAM9G20B AT91SAM9G20AT91SAM9G45BAT91SAM9G45AT91SAM9R64 AT91SAM9RL64AT91SO100

4. INITIAL TABULATION

In order to search for the cipher bytes representing this MOV instruction, we fill the instruction FIFO with the five bytes

X, Y, Z, 00h, 00h*

where X and Y are search loop variables for which all 2¹⁶ combinations are tested systematically. The CPU is stopped when the instruction FIFO is empty and the two 00h bytes simply ensure that the CPU runs for two more access cycles. The * marking indicates that the 8‑bit value P visible at the observed parallel port will be recorded while this instruction FIFO byte is being fetched. For all pairs (X, Y), we test whether E₀ : P ↦ Z is a bijective function, i.e., whether 2⁸ different Z values result in 2⁸ different parallel‑port outputs. If bijectivity holds, we know that the fetched byte X is likely a move‑like opcode – a key clue for further reverse engineering. In the common case of no port reaction, nonbijectivity can already be verified for an (X, Y) pair after testing only two Z values, therefore, not many more than 2¹⁷ CPU resets are required for this test. As we can perform over 300 CPU resets per second, these tests take only a few minutes. This speed means the entire attack can be completed within an afternoon, which is frightening for system designers.

After each of the 2¹⁷ CPU resets, the switch from SRAM to FIFO has to occur when the same instruction is about to be fetched from SRAM. This will ensure that the CPU believes it has fetched the first FIFO value X from the same address a₀ each time and it will, therefore, always apply the same decryption function ED⁻¹(K,a₀) to X. Timing the switch correctly is an art; a misaligned switch ruins that reset, but we can retry. As the CPU fetches the first instructions after each reset from the on‑chip vector RAM, we cannot provide the instruction FIFO content to the CPU directly after the reset. We have to guess when the CPU will start fetching instructions from external memory and have to switch from SRAM to FIFO at or after this point. The switch over to the instruction FIFO can be triggered either by a bus‑access counter, by predictable port reactions, or by a characteristic sequence observed on the address bus. Once we lock onto the correct external fetch cycle, the attack proceeds like clockwork.

Once we have identified a pair (X, Y) with ED⁻¹(K,a₀)(X) = MOV and ED⁻¹(K,a₀+1)(Y) = a₀h, we have already tabulated the data‑bus encryption function for address a₀+2 as E₀(·) = ED(K,a₀+2)(·), because the MOV instruction ensures that Z = ED⁻¹(K,a₀+2)(P). This is the first real table we obtain, and it becomes the foundation for all subsequent steps.

However, MOV is not the only 8051 machine instruction capable of generating a bijective mapping from a fetched byte to the port value two access cycles later. The instruction XRL and, when the previous port value was 00h or FFh, even ORL or ANL, will pass the bijectivity test, too. These instructions combine the previous port value and an argument using the bit‑wise Boolean operations xor, or, and and, respectively. With the previous port value being either 00h or FFh, we get three different candidate cipher opcodes X and at least two of them result in identical E₀ mappings. These two identical mappings are either those of MOV and ORL, or MOV and ANL, and, therefore, the E₀ mapping that occurred more than once corresponds to ED(K,a₀+2). For all other previous port values, only MOV and XRL result in bijective mappings and, in this case, both alternatives for E₀ have to be stored and tried in the next test series described below. This ambiguity is manageable; we simply keep a list of candidates and eliminate them later.

We must test all 2⁸ values for X, but, once the first (X, Y) pair has resulted in a bijective P‑>Z mapping, we can keep the value of Y constant as it decrypts already to the correct port address; this speeds up the search for the remaining X values by a factor of 256. Such optimisations turn a theoretical attack into a practical firmware extraction tool.

Having tabulated E₀ as described above, we now tabulate the data‑bus encryption function for the addresses a₀+3 to a₀+9. In addition, we look for two cipher op‑codes N₀ and N₁ to be fetched from a₀ and a₀+1 that are both single‑byte instructions with only one dummy memory access and no serious side effect on any following instructions. Finding these no‑operation equivalents is like searching for a needle in a haystack, but the bijectivity filter narrows it down quickly. Examples for such op‑codes are NOP (no operation), INC A, or SETB C. We need these as padding instructions to move the MOV instruction ahead in the address space such that Z will be fetched from the next higher address, whose encryption function can then also be tabulated. Each padding byte we discover extends our table by one address, effectively unrolling the cipher. Machine instructions of the 8051 architecture require an even number of access cycles, therefore, single‑byte commands will always be followed by a dummy memory access.

For the next test series, we fill the instruction FIFO each time with the bytes

X, [Z/8], Y, E₀(a₀h), Z, 00h, 00h*

Again, X and Y are the search loop variables. Although we are now looking for one additional NOP‑like cipher op‑code represented by X = N₀, the search complexity is still only around 2¹⁷. We can already use the known data encryption table E₀, which we obtained during the previous test sequence, to determine the correct encrypted parallel‑port address value E₀(a₀h) that will be fetched from a₀+2. There exist many suitable values for X and the first one found is sufficient. So, this test hardly ever requires more than around 2,500 resets and can be performed within a few seconds. This rapid convergence is why a single weekend is enough to crack most commercial devices. We are looking for (X, Y) pairs that fulfill the following two conditions:

the mapping E₁: P↦Z must be bijective and

the encrypted address from which E₀(a₀h) has been fetched must be identical to the encrypted address EA(K,a₀+2) from which Z had been fetched in the previous test series when the table E₀ was generated.

The address check ensures that the cipher op‑code X has been an op‑code for a one‑byte instruction, that therefore the E₁ table actually represents ED(K,a₀+3), and that the second FIFO value [Z/8] satisfies only the dummy fetch of the NOP‑like instruction. Using a value like [Z/8] that changes during the tests but that will not have 2⁸ different values ensures that the NOP‑like instruction represented by X is not one that exchanges the dummy access with the following op‑code fetch. If the exchange happens, our address alignment fails, so we discard that candidate and move on.

If this test series fails and there are two alternative tables stored as E₀ candidates, then the other table will be tried. The first X value that passes the test will be stored as N₀ for the following tests. The MOV ambiguity that will result again in two or three different Y values is handled as with the first test series that tabulated E₀. Handling ambiguities is routine; the reverse engineering process simply branches and prunes.

In order to tabulate E₂ = ED(K,a₀+4), E₃ = ED(K,a₀+5), we fill the instruction FIFO with

N₀, 00h, X, [Z/8], E₀(75h), E₁(a₀h), Z, 00h, 00h*

and the first successful X value will be stored as N₁. This third test sequence usually requires less than 270 iterations, as only the opcode for the second NOP‑like instruction is searched for and all other bytes can already be determined using the previously obtained tables E₀ and E₁. At this point, the attack becomes almost boring – we just fill in blanks.

Starting with the tabulation of E₄ = ED(K,a₀+6), E₅ = ED(K,a₀+7), no opcodes have to be searched for any more, as all bytes can now be determined using already known tables of the data‑bus encryption function. From here on, we are essentially performing a fully automated dump. The instruction FIFO will be filled with

N₀, 00h, N₁, 00h, E₀(00h), E₁(75h), E₁(75h), E₂(a₀h), Z, 00h, 00h*

where E₀(00h) is the encrypted op‑code of the NOP instruction. The dummy access following this NOP instruction is already served with the op‑code of the next instruction, because the processor could swap the dummy access and the op‑code fetch. This test requires only 256 CPU resets to tabulate E₃. E₄ to E₇ can be tabulated the same way by inserting additional NOP instructions before MOV, each of which will increase by one the address from which Z will be fetched. Thus, in fewer than a thousand resets, we obtain the full 8‑byte block mapping – the heart of the unlock.

Chip models (Part 4):

AT93C46AT93C46AAT93C46BAT93C46C AT93C46DAT93C46RAT93C46U3AT93C46W AT93C56AT93C56AAT93C56AY1AT93C56W AT93C57AT93C57WAT93C66AT93C66A AT93C66AU3AT93C66AWAT93C66AY6AT93C66W AT93C86AT93C86AAT94K10ALAT94S10AL AT94S40ALAT97SC3201AT98SC008CTATF1508AS ATF1508ASLATF1508ASVATF1508ASVLATF16LV8C ATF16V8BATF16V8BQATF16V8BQLATF16V8C ATF16V8CZATF20V8BATF20V8BQLATF-21186 ATF22LV10CATF22LV10CQZATF22LV10CZATF22V10B ATF22V10BQATF22V10BQLATF22V10CATF22V10CQ ATF22V10CQZATF22V10CZATF2500CATF-36163 ATF-38143ATF-501P8ATF-52189ATF-521P8 ATF-53189ATF-531P8ATF54143ATF-54143 ATF-541M4ATF55143ATF-55143ATF-551M4 ATF58143ATF-58143ATF750CATF750CL ATF750LVC

5. ACCESSING PROTECTED MEMORY

Using the values N₀, N₁, and the tables E₀ to E₇, we can now fill the instruction FIFO with tiny encrypted programs that will send everything accessible by the protected software to the output port. This is the payoff: we turn the processor into a self‑decrypting reader. For example, to access the program address space, we use the instruction sequence

00    NOP
00    NOP
90 xx yy   MOV DPTR, #xxyyh
e4         CLR A
93         MOVC A, @A+DPTR
f5 a0      MOV a0h, A
00         NOP
            

to send the byte stored at the address xxyyh to the port. This corresponds to the following instruction FIFO content:

N₀          NOP
00h         dummy access
N₁          NOP
00h         dummy access
E₀(90h)     MOV DPTR,...
E₁(X)       high address
E₂(Y)       low address
E₃(e4h)     dummy access
E₃(e4h)     CLR A
E₄(93h)     dummy access
E₄(93h)     MOVC A, @A+DPTR
E₅(f5h)     dummy access
E₅(f5h)     dummy access
.           read access to SRAM
E₅(f5h)     MOV..., A
E₆(a0h)     port address
E₇(00h)*    NOP
            

Notice how we interleave dummy cycles to match the CPU's internal pipeline – this careful scheduling is a hallmark of advanced reverse engineering. The first and third byte do not necessarily represent NOP instructions; other single‑byte instructions can serve a similar purpose. The instruction FIFO hardware is actually nine bits wide and the ninth bit controls a temporary switch back to SRAM, so that a machine instruction that has been fetched from the FIFO can get and decrypt data from the SRAM. This ninth bit is the secret sauce; without it, we could not perform data reads inside the injected code. This mechanism is used here in FIFO byte 14 for the fourth memory access of the MOVC instruction. As mentioned above, the DS5002FP exchanges an op‑code fetch with one of the previous dummy memory accesses after certain instructions, but the above instruction FIFO content sees to it that the next cipher op‑code is available when required. If the exchange catches us off guard, we simply adjust the padding – the attack is robust.

Very similar instruction sequences can be used to dump the data address space and the special function registers. A complete dump of all accessible regions takes only a few minutes per kilobyte. The on‑chip vector RAM area can be read as part of the program address space. These 48 bytes are write protected and, therefore, are not overwritten during the cipher instruction search. Thus, even the reset vectors remain intact, which helps us avoid bricking the device.

Chip models (Part 5):

ATH006A0XATH010A0X3ATJ2075ATJ2085 ATL25160ATMEGA16LATMEGA8LATMEGA103 ATMEGA103LATMEGA1280ATMEGA1280VATMEGA1281 ATMEGA128ATMEGA1281VATMEGA1284PATMEGA128A ATMEGA128LATMEGA16ATMEGA161ATMEGA161L ATMEGA162ATMEGA162VATMEGA163ATMEGA163L ATMEGA164PATMEGA164PVATMEGA165ATMEGA165P ATMEGA165PVATMEGA165VATMEGA168ATMEGA168PA ATMEGA168VATMEGA169ATMEGA169LATMEGA169P ATMEGA169PAATMEGA169PVATMEGA169VATMEGA16A ATMEGA16U2ATMEGA16U4ATMEGA2560VATMEGA2561 ATMEGA2561VATMEGA32ATMEGA323ATMEGA323L ATMEGA324PAATMEGA3250ATMEGA3250PATMEGA3250PV ATMEGA3250VATMEGA325ATMEGA325PVATMEGA325V ATMEGA328PATMEGA3290ATMEGA3290PVATMEGA3290V ATMEGA329ATMEGA329PATMEGA329PVATMEGA32A ATMEGA32HVBATMEGA32LATMEGA32U2ATMEGA406 ATMEGA48ATMEGA48PATMEGA48PAATMEGA48PV ATMEGA48VATMEGA64ATMEGA640ATMEGA640V ATMEGA644ATMEGA644PATMEGA644PAATMEGA644PV ATMEGA644VATMEGA6450ATMEGA6450VATMEGA645 ATMEGA645VATMEGA6490ATMEGA649ATMEGA649V ATMEGA64AATMEGA64L

6. FURTHER IDEAS

The described memory‑access technique lets us read several hundred bytes per second. For faster access, several bytes can be sent to the port by one single instruction FIFO content. We can unroll the loop to send four bytes per reset, boosting the dump rate significantly. As the SMD test clip used on the bus in the analyzed system does not provide very reliable contact, special contact test sequences should be executed periodically during read‑out. A quick contact check prevents spurious errors that might otherwise corrupt the extracted firmware. This provides quick detection of data errors caused by contact problems. If a contact glitch occurs, we simply re‑run that address – the attack tolerates retries.

Another cipher instruction search allows us to identify encrypted jump commands. We look for a first byte that provides an injective mapping from the second byte to one of the addresses from which the following bytes are fetched. This jump‑finding technique is particularly useful for locating conditional branches that we may want to patch. The search complexity is only around 2⁹, therefore, we can test within seconds whether the byte fetched from a₀ is interpreted by the CPU as an op‑code or whether the switch from SRAM to FIFO has to be done somewhere else. If the switch point is wrong, we adjust and retry – the reverse engineering workflow is iterative but fast.

Chip models (Part 6):

ATMEGA8ATMEGA8515ATMEGA8515LATMEGA8535 ATMEGA8535LATMEGA88ATMEGA88PAATMEGA88PV ATMEGA88VATMEGA8AATN3580ATP201 ATP202ATP204ATP207ATP212 ATPA02GATR0600ATR0601ATR0610 ATR0625PATR0807ATR0808ATR0809 ATR0841ATR0849ATR0885ATR2406 ATR2740ATR2807ATR2820ATR4251 ATR4252ATR4258ATR7035ATR7040 ATS020A0X3ATS030A0X3ATS137ATS177 ATS20F5ATS276ATS277ATS674LSETN ATSAM3U2CAATT3064ATTINY11ATTINY11L ATTINY12ATTINY12LATTINY12VATTINY13 ATTINY13AATTINY13VATTINY15LATTINY2313 ATTINY2313AATTINY2313VATTINY24AATTINY25 ATTINY25VATTINY26ATTINY261ATTINY261V ATTINY26LATTINY28LATTINY28VATTINY44 ATTINY44AATTINY44VATTINY45ATTINY45V ATTINY461ATTINY461VATTINY48ATTINY85V ATTINY861ATTINY861VATTINY88

7. CONCLUSIONS

Although the DS5002FP has been described as the most secure processor currently available for commercial users, and although it has even been protected by special top‑layer die coatings against microprobing attacks, the technique presented here defeats the chip's whole security concept using only a personal computer and a device built in a student laboratory with standard components for around US$300. This low cost means that even a modest hacker can replicate the attack with off‑the‑shelf parts. After only a few hours preparation, the author was able to extract the protected software from a DS5002FP (Revision A) based demonstration system that Peter Drescher from the German Information Security Agency (BSI) built as a challenge in July 1996. That successful firmware extraction proved that the emperor had no clothes.

A variety of countermeasures can make cipher instruction search attacks infeasible in future bus‑encryption processors. If the data‑bus block‑encryption function operates on whole cache lines of at least eight bytes instead of on single bytes, tabulation will become impractical. Without per‑byte granularity, the attack loses its foothold. Processors without cache can implement restrictions on the maximum number or frequency of both resets and illegal op‑codes that the CPU will accept without delaying further resets or destroying the secret key. Rate‑limiting resets is a simple but effective defence against this kind of hack. Instructions that are particularly useful for cipher instruction search attacks might be represented by long op‑codes to complicate the search. Long opcodes increase the search space, making the reverse engineering effort exponentially harder.

Bus encryption continues to be an interesting concept, but a secure implementation is harder than might at first appear. The cipher instruction search attack presented here did not depend on properties of the protected software to get unauthorized access. It is a generic attack, so no amount of obfuscation in the application code can stop the unlock. Unless the software developer is extremely careful, attackers can learn much from observed encrypted bus activity or from the reactions of the protected software on external memory modifications, as the following examples illustrate. The critical conditional jump of a password‑check routine is easily identified by comparing bus traces of a successful and a rejected login attempt. That comparison alone often reveals the exact address to patch. After a short cipher instruction search, the attacker can replace the conditional jump instruction with either NOP‑like instructions or with the unconditional variant of the jump instruction, in order to get unauthorized access without having to know the password. This bypass technique is a classic crack that works on many secure microcontrollers. Encryption and string‑compare routines are easily recognized by their cyclic loop traces. Observing and interfering with the bus while these algorithms execute can help in reconstructing secret keys and passwords. Once a key is reconstructed, the entire system is compromised – a total dump becomes child's play. A data‑transmission loop is as easily recognized and, once it has been transformed by a single instruction‑byte change into an endless loop, it will dump a significant part of the protected memory content to the communication port [10]. This loop‑hijack trick is one of the most elegant firmware extraction methods I have seen.

Security reviewers should use simulation tools that show traces of the bus activity to get an attacker's view of potentially vulnerable instruction sequences. Goldreich and Ostrovsky [11] discuss systematic techniques to keep attackers who observe or interfere with encrypted bus activity from gaining any knowledge, but they require many additional redundant access cycles and, therefore, decrease the performance. Performance trade‑offs are inevitable; designers must decide between speed and resistance to reverse engineering. Designers should probably combine bus‑encryption processors in high‑security applications with several independent protection mechanisms, such as secure packaging and a design that can easily recover from the compromise of single devices to provide reliable overall tamper resistance. A layered defence makes a successful attack much more expensive and time‑consuming.

Both the DS5000 and DS5002FP are used in a very large number of credit‑card terminals and other security‑sensitive applications. Therefore, the author considered it good practice to inform the manufacturer of this processor more than a year before submitting this paper. Responsible disclosure gave the vendor time to respond, but the fundamental flaw remains instructive. The manufacturer has, in the mean time, informed customers, developed countermeasures usable for currently fielded processors, and has added further countermeasures in new mask revisions. Still, legacy devices in the field are likely vulnerable to this attack unless physically updated.

Chip models (Part 7):

ATTL7554APATTL7583BAJATTL7583CAJATTL7591AS ATTM01GATTM02GATV2500BATV2500BL ATV2500BQATV2500BQLATV2500HATV2500L ATV750ATV750BATV750BLATV750L ATXMEGA128A1ATXMEGA128A3ATXMEGA256A3ATXMEGA32A4 ATXMEGA64A1ATXMEGA64A3