Reverse engineering is a foundational and critical discipline within the field of software security. It is the sophisticated process of deconstructing compiled software—the opaque binary—to unveil its internal logic, design, and implementation details. In environments where the original source code is intentionally unavailable, such as with proprietary applications, or lost due to data degradation, this practice serves as the indispensable bridge between low-level machine instructions and human-readable understanding. 

For software security professionals, reverse engineering is not merely an academic exercise; it is an essential toolkit for proactive defense and analysis. Its core applications include deep-dive malware analysis to understand attacker techniques, vulnerability research to discover and patch zero-day flaws, maintenance of mission-critical legacy systems, and ensuring legitimate software interoperability and compliance. 

This article offers a comprehensive exploration of the core techniques employed to dissect these binaries, progressing from initial surface-level assessment and rigorous static examination to controlled dynamic execution and the frontier of advanced, automated reasoning and decompilation.

1. strings

The strings utility is a command-line tool designed to extract printable character sequences from binary files. While compiled executables consist primarily of machine code unreadable by humans, they frequently contain embedded text. This utility acts as a filter, scanning non-text files to locate and display coherent ASCII or Unicode sequences. It operates by parsing a file to identify chains of characters—such as letters, numbers, and punctuation—that meet a minimum length requirement. This process allows for the retrieval of readable data without the need for source code or program execution.

The output often provides insight into a program’s internal configuration and dependencies. Common artifacts revealed include hardcoded file system paths pointing to configuration files, logs, or libraries, as well as network indicators like embedded IP addresses, URLs, and API endpoints used for communication. Additionally, the tool exposes program logic through error messages and status updates that clarify how exception states are handled, while metadata such as copyright notices and version information help establish software provenance.

In security contexts, strings serves a critical role in examining suspicious binaries. It facilitates the discovery of Indicators of Compromise (IOCs), such as Command and Control (C2) domains or malicious IP addresses. Furthermore, the extracted text can reveal intended malicious behavior, such as specific system commands or registry modifications. Conversely, a lack of readable output—or output consisting entirely of random characters—may indicate that a binary has been packed or obfuscated to conceal its content.

While this level can provide valuable information about the intended behaviour of a program, is it usually not enough. So, in order to better understand the program’s actual structure and behaviour, another technique can be used: static analysis.

strings in linux
strings in linux
strings in Linux

2. static analysis

Static analysis involves the examination of software code or compiled binaries without executing the program. In the field of reverse engineering, this methodology provides the means to reconstruct the logic, architecture, and design of a software application when the original source code is unavailable.

This approach is primarily used to understand “black box” binaries. By treating the compiled file as a fixed dataset, it becomes possible to recover algorithms, understand data structures, and verify how a compiler has translated high-level instructions into machine code. This process is essential for achieving interoperability, as it enables the comprehension of undocumented file formats or network protocols required for system compatibility. Additionally, it supports legacy maintenance by recovering lost logic from older software versions lacking documentation. Finally, the methodology allows for logic recovery, where mapping control flow—such as loops, branches, and function calls—helps reconstruct the original programming intent.

The process relies on advanced software suites capable of translating machine code back into human-readable assembly language or pseudo-code:

  • IDA (Interactive Disassembler): A widely used platform known for its ability to create interactive control flow graphs. It visualizes the binary’s structure and facilitates navigation through complex function hierarchies.
  • Ghidra: An open-source software reverse engineering suite that includes a highly capable decompiler. It allows for the simultaneous analysis of a binary by multiple users and supports scripting to automate repetitive tasks.
  • RetDec: A retargetable decompiler used to convert machine code from various architectures back into high-level C code. It is often employed to cross-reference or verify the output generated by other decompilers.

For tasks requiring less overhead than a full graphical suite, objdump serves as a lightweight, command-line utility in Linux. It allows for the rapid inspection of a binary’s low-level details, such as file headers, symbol tables, and memory section layouts. It is frequently used to quickly verify architecture information or to generate a raw linear listing of the assembly instructions contained within a file.

Static analysis provides a detailed blueprint of the code, but that blueprint doesn’t always reveal how the program behaves in a live environment. To observe the program’s real time behaviour, dynamic analysis can be used.

ghidra in linux
ghidra in linux
Ghidra in Linux

3. dynamic analysis

Dynamic analysis involves evaluating a software program by executing it and observing its behavior in real-time. This approach stands in contrast to static analysis, which examines the code as a fixed blueprint without running it. While static methods focus on the theoretical structure and potential logic of a binary, dynamic analysis observes the program as a functioning machine. This distinction is crucial because running the code reveals the “true” state of the program, including values that are calculated only at runtime, such as decrypted strings or dynamically loaded libraries, which remain invisible to a static examination.

In the Linux and macOS ecosystems, tools like gdb (GNU Debugger) and lldb serve as the standard instruments for this task. Although frequently employed for general software debugging, in a reverse engineering context, they function as precision tools for dissecting compiled binaries. These debuggers provide granular control over execution, allowing the program to be paused at specific memory addresses or functions using breakpoints. Once execution is suspended, the contents of CPU registers and system memory can be inspected directly. This capability exposes the actual arguments being passed to functions, bypassing the need to mentally simulate complex data transformations. Furthermore, these tools allow for instruction stepping, where the program is advanced one instruction at a time to precisely trace logic flow through complex branches.

This level of control becomes particularly useful in scenarios where static analysis reaches a dead end. For instance, many binaries are “packed” or encrypted to hide their code; dynamic analysis allows the program to run naturally until it decrypts itself in memory, at which point the unpacked code can be captured. It also serves as a method for logic verification, where hypotheses formed during static analysis can be tested against reality. If a specific function is believed to generate a cryptographic key, a debugger can confirm the output values in real-time. Finally, these tools are essential for fault analysis; if a specific input causes a program to crash or behave unexpectedly, the debugger can isolate exactly which instruction fails, providing deep insight into internal error handling and stability. Even with dynamic analysis, manually exploring every possible branch of a complex program can be incredibly time-consuming. To automate the discovery of execution paths and pinpoint the exact conditions that trigger specific behaviours, symbolic execution can be used.

4. symbolic execution

Symbolic execution is an automated analysis technique that explores a binary’s logic by treating inputs as algebraic variables rather than concrete values. Instead of executing a program with a specific string or number, the analysis engine processes symbols. When the execution encounters a branching decision, such as an if-else statement, it “forks” into parallel paths, mathematically tracking the constraints required to navigate each route.

The primary strength of this methodology is its ability to mathematically derive the precise input needed to reach a specific code location. By accumulating path constraints and processing them through a Satisfiability Modulo Theories (SMT) solver, the technique can generate concrete values that satisfy complex logic checks. This makes it invaluable for bypassing software protection schemes, generating valid serial keys, or discovering the exact input required to trigger a deep-seated vulnerability.

In the reverse engineering landscape, angr is the most prominent Python-based framework for this task. It allows users to target specific functions or code blocks, effectively instructing the engine to “find the input that reaches this address.” However, due to the computational cost of tracking exponential path combinations—a problem known as “path explosion”—angr is most effective when applied to small, mathematically complex components rather than entire applications.

Several platforms facilitate this advanced analysis:

  • angr: A versatile, Python-centric framework widely used for binary analysis and solving CTF challenges.
  • Triton: A dynamic binary analysis library offering customizable components for symbolic execution and taint analysis.
  • Manticore: A tool developed by Trail of Bits, capable of analyzing Linux ELF binaries and Ethereum smart contracts.
  • KLEE: A symbolic virtual machine built on the LLVM infrastructure, highly effective for high-coverage bug finding.

conclusion

Unraveling compiled code requires a layered approach. The process typically begins with surface inspection and structural mapping, then evolves into observing the program’s behavior in real time. When manual investigation reaches its limit, automated mathematical reasoning provides the precision to solve complex logic. Ultimately, the ability to synthesize these diverse perspectives is what transforms opaque machine instructions into a transparent and intelligible system.

references:

about the author.
Sandu G
Sandu G

Sandu G

software engineer, randstad digital

Sandu started his journey with Randstad Digital in 2021.
He is passionate about technology in general. His area of expertise is software development, with a special interest in software security.