Skip to content

Malware Analysis Tools Proficient

🔒 Security & DevOps
⏱️ ~4 days 📚 Prerequisites: File Handling, Automation

When you'd use this

Using Python to analyze and detect malicious software — defensively.

Analyze suspicious binaries and behavior safely for defensive security research.

What you'll learn

  • What defensive malware analysis is
  • Static analysis: hashing and indicators (tested)
  • Dynamic analysis in sandboxes
  • Automating analysis workflows
  • The tools of the trade

Malware analysis is the defensive discipline of examining malicious software to understand what it does, detect it, and protect against it. This is what security researchers, incident responders, and antivirus teams do. Python is the field's dominant scripting language. The hashing/indicator example here is run-verified.

This page is strictly defensive

This covers analyzing and detecting malware to protect systems — the work of blue teams, researchers, and responders. It does not cover writing malware. Always analyze suspected malware in an isolated, disposable environment (a VM with no network or a dedicated sandbox), never on a machine you care about or a production network.


Two kinds of analysis

Static (inspect without running) vs dynamic (observe behavior in a sandbox).

   STATIC analysis                 DYNAMIC analysis
   examine WITHOUT running          run in an isolated sandbox
   (hashes, strings, structure)     and observe behavior
   safe, fast, but limited          reveals runtime behavior, riskier
  • Static — inspect the file without executing it: compute hashes, extract strings, examine structure. Safe and quick.
  • Dynamic — run the sample in a controlled, isolated sandbox and watch what it does (files it touches, network it contacts). More revealing, but must be tightly contained.

Static analysis: hashing & indicators (tested)

Static analysis: hashing & indicators in Malware Analysis Tools — what it is and when to use it.

The first step in analyzing any suspicious file: compute its hash (to identify it and check against known-malware databases) and scan for indicators. Runnable:

import hashlib

def file_hashes(data: bytes) -> dict:
    """Compute standard hashes used to identify a sample."""
    return {
        "md5": hashlib.md5(data).hexdigest(),
        "sha256": hashlib.sha256(data).hexdigest(),
    }

def find_indicators(data: bytes) -> list[str]:
    """Flag suspicious strings (indicators of compromise)."""
    text = data.decode("latin-1", errors="ignore")
    suspicious = ["cmd.exe", "powershell", "http://", "CreateRemoteThread"]
    return [s for s in suspicious if s in text]

sample = b"harmless file that mentions http:// and cmd.exe somewhere"
print(file_hashes(sample)["sha256"][:16], "...")
print("indicators:", find_indicators(sample))

Output:

c0dd4362e097... ...
indicators: ['cmd.exe', 'http://']

The hash uniquely identifies the file — you'd look it up on VirusTotal or a threat-intel feed to see if it's known malware. The indicator scan flags suspicious strings (references to shells, URLs, injection APIs) that warrant deeper analysis. This is exactly how basic triage tools and YARA rules start — pattern-matching on file contents.


Dynamic analysis: sandboxes

Run suspicious code in isolation and watch what it does.

To see what a sample actually does, run it in an isolated sandbox that records its behavior — files created, registry/config changes, network connections, processes spawned. Because you're running real malware, isolation is non-negotiable:

  • Cuckoo Sandbox — the classic open-source automated malware sandbox (Python-based).
  • Cloud sandboxes — Any.run, Hybrid Analysis, VirusTotal's behavioral analysis.
  • Isolated VMs — a disposable VM snapshot with no network (or a controlled fake network), reverted after each run.

The output is a behavior report: "this sample downloaded X, wrote Y, contacted Z" — the basis for detection signatures and incident response.


Automating analysis

Script triage so you can process many samples quickly.

Python excels at gluing analysis workflows together (see Automation & Scripting):

# Documented workflow shape (uses external services/libs)
# 1. Hash the sample
# 2. Query threat intel (e.g. VirusTotal API) by hash
# 3. Extract static features (strings, structure)
# 4. If unknown, submit to a sandbox for dynamic analysis
# 5. Aggregate into a report; generate detection rules (YARA)

Common Python libraries in the field (documented; not installed here): pefile (parse Windows PE executables), yara-python (pattern-matching rules), capstone (disassembly), oletools (malicious Office docs), plus API clients for threat-intel services.

Analysis libraries follow documented usage

pefile, yara-python, capstone, and threat-intel APIs aren't installed here — the hashing/indicator example is run-verified. These libraries do the heavy lifting of parsing executable formats, matching signatures, and disassembling code.


The tools of the trade

YARA, disassemblers, and sandboxes defenders rely on.

Task Tool
Identify by hash hashlib + VirusTotal API
Pattern/signature matching YARA (yara-python)
Parse Windows executables pefile
Disassembly capstone, radare2, Ghidra
Malicious documents oletools
Automated dynamic analysis Cuckoo Sandbox
Kernel/behavior tracing see Kernel Tracing

The defender's mindset

Analyze to protect — contain samples and never run them unsafely.

Malware analysis feeds the defensive cycle: analyze a sample → understand its behavior → write detection (YARA rules, IDS signatures) → deploy protection → respond to incidents. It connects to Vulnerability Scanning (what weaknesses malware exploits) and Kernel Tracing (observing behavior at the OS level).

Contain everything

Never execute a suspected sample outside strict isolation. Use a VM with networking disabled (or a controlled analysis network), take a snapshot beforehand, and revert afterward. Treat every sample as live and dangerous — because it is.


Practice exercises

  1. Extend find_indicators with more indicator strings and return the byte offset where each was found.
  2. Compute and compare hashes of two files to detect if they're identical (deduplication/triage).
  3. Explain the difference between static and dynamic analysis and when you'd use each.
  4. Describe why dynamic analysis must run in an isolated, revertible environment.
  5. Sketch (in comments/pseudocode) an automated triage pipeline: hash → threat-intel lookup → indicator scan → report.

💬 Discussion

Have a question about this topic? Found an error? Share your thoughts below.