Skip to content
AsteriaX Labs

Research direction 02

Malware Detection & Deep Learning

Malware detection that combines deep representation learning with heuristic search for feature selection, starting from careful replication of existing work.

Status
Baseline Investigation
Lead
Yoga
Current stage
Studying and replicating an external published approach.

Context

Malware detection systems work with high-dimensional, noisy features extracted from executables and their behaviour. Which features a model sees, and how they are represented, shapes what it can learn.

Deep representation learning can compress raw features into more informative representations. Heuristic search algorithms can explore the space of feature subsets that would be impractical to search exhaustively. Hybrid approaches combine the two.

Before proposing anything new, this direction starts by understanding and replicating an existing hybrid approach. A reliable baseline is a precondition for any later claim of improvement.

Research areas

Malware Detection
Classifying software samples as malicious or benign from extracted static or behavioural features.
Deep Representation Learning
Learning compact representations of input features with neural networks.
Feature Selection
Choosing a subset of features that preserves useful signal while reducing dimensionality.
Heuristic Search
Search strategies that explore large combinatorial spaces, such as feature subsets, without exhaustive enumeration.
Reproducible Experiments
Fixed seeds, documented configurations, and recorded deviations, so results can be checked and repeated.

Current focus

Baseline replication

The research foundation is an external paper: “Improving malware detection performance using hybrid deep representation learning with heuristic search algorithms.” It is under investigation for baseline replication. It was not authored by Yoga or any AsteriaX Labs collaborator.

The aim at this stage is to reproduce the described approach as faithfully as the publication allows, and to document every point where the description leaves implementation choices open.

Thesis novelty has not been defined. That decision depends on what the replication shows.

Fig. 02 · Hybrid detection pipeline

01SamplesRaw features02RepresentationDeep learning03SelectionHeuristic search04ClassifierMalicious / benignCandidate feature subsetsSchematic only. Configuration is under replication.
Schematic of the general hybrid approach: learned representations, heuristic search over feature subsets, then classification. Exact configuration is the subject of replication. Not a result.

Questions guiding the replication

  1. Q1

    Can the described pipeline be reproduced from the published information alone?

  2. Q2

    Which implementation details are underspecified, and how much do reasonable choices for them matter?

  3. Q3

    How stable are selected feature subsets across random seeds and data splits?

  4. Q4

    What does a well-defined baseline need to include so that later comparisons are meaningful?

Replication plan

A plan of work, not a record of completed steps.

  1. Study the publication

    Extract the pipeline, data handling, and evaluation protocol, and list every unstated assumption.

  2. Prepare data

    Assemble the dataset and preprocessing as closely to the paper’s description as possible.

  3. Implement the representation stage

    Build the deep representation learning component and verify its behaviour in isolation.

  4. Implement heuristic feature selection

    Implement the search procedure over feature subsets and record its configuration.

  5. Evaluate and document

    Run the documented protocol across seeds and record agreements and discrepancies with the publication.

Scope and claims

What this work investigates

  • Whether an existing hybrid approach can be reproduced reliably.
  • Sensitivity of the approach to implementation choices the publication does not fix.
  • What a sound, reusable baseline for later work looks like.

What it does not claim

  • Authorship of the paper under investigation.
  • A defined thesis contribution or novelty. This is not yet decided.
  • Benchmark results or completed experiments.
Malware DetectionDeep Representation LearningFeature SelectionHeuristic SearchReproducible Experiments