Sources

Overview

Two application threads share a computational text and network toolkit: (1) comparing media content similarity across news agencies, and (2) extracting relational structure from narrative text (film characters) and structured HR records.

Evidence

  • News agency text exhibits measurable similarity clusters by outlet and topic
  • Character dialogue and scene structure yield analyzable interaction networks
  • HR tabular data supports predictive features orthogonal to text pipelines
  • Methods overlap in preprocessing discipline but differ in representation (vectors vs graphs)

Analysis

Text mining projects in this portfolio are method demonstrations tied to substantive domains — media systems, cultural narrative, labor data — rather than generic NLP benchmarks. The unifying skill is translating messy symbolic data into matrices or graphs suitable for statistics and machine learning.

For Suengj Note, keeping these threads in one research entry clarifies the shared tooling without merging incompatible substantive claims.

References

Related