API reference¶
SkillExtractorRefactored is the public entry point. Construct it once with a model choice, then call
extract_concepts on as many tables as you like; the model and taxonomy indexes are loaded once and
reused.
This page is generated from the docstrings in
laiser/skill_extractor_refactored.py.
SkillExtractorRefactored ¶
SkillExtractorRefactored(model_id: Optional[str] = None, hf_token: Optional[str] = None, api_key: Optional[str] = None, use_gpu: Optional[bool] = None, backend: Optional[str] = None, temperature: float = DEFAULT_TEMPERATURE, seed: Optional[int] = GENERATION_SEED)
Refactored skill extractor with improved separation of concerns.
This class provides a clean interface while delegating specific responsibilities to appropriate service classes.
Initialize the skill extractor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_id
|
str
|
Model ID for the LLM |
None
|
hf_token
|
str
|
HuggingFace token for accessing gated repositories |
None
|
api_key
|
str
|
API key for external services (e.g., Gemini) |
None
|
use_gpu
|
bool
|
Whether to use GPU for model inference |
None
|
backend
|
str
|
Backend to use for LLM inference (e.g., "llama_cpp", "huggingface", "openai", "gemini") |
None
|
temperature
|
float
|
Decoding temperature for whichever backend is used. Defaults to 0.0 (greedy decoding), so repeated runs over the same input produce the same extraction. Raise it only when varied output is wanted. |
DEFAULT_TEMPERATURE
|
seed
|
int
|
Seed for backends that accept one: vLLM, llama.cpp, Gemini, and local Transformers models when temperature is above 0. Defaults to 42; pass None to leave it unset. The OpenAI API accepts no seed, so there reproducibility rests on greedy decoding alone. |
GENERATION_SEED
|
extract_concepts ¶
extract_concepts(data: DataFrame, id_column: str = 'Research ID', text_columns: List[str] = None, input_type: str = 'job_desc', top_k: Optional[int] = None, similarity_threshold: Optional[float] = None, levels: bool = False, batch_size: int = DEFAULT_BATCH_SIZE, warnings: bool = False, allowed_sources: Optional[List[str]] = None, concepts: List[str] = None, extract: List[str] = None, return_edges: bool = False, similarity_thresholds: Optional[Dict[str, float]] = None, timing: bool = False, output_csv_path: Optional[str] = None)
Extract concepts from text and align them to taxonomy data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
DataFrame
|
Input dataset |
required |
id_column
|
str
|
Column name for document IDs |
'Research ID'
|
text_columns
|
List[str]
|
Column names containing text data |
None
|
input_type
|
str
|
Type of input data |
'job_desc'
|
top_k
|
int
|
Maximum number of aligned matches per concept type, per document (default: 25). With skills, knowledge and tasks all requested, one document can return up to three times this many rows. |
None
|
similarity_threshold
|
float
|
One minimum similarity score applied to every type. When omitted, the per-type defaults below apply. |
None
|
similarity_thresholds
|
dict
|
Per-type minimums, keys "skill", "knowledge", "task". Overrides similarity_threshold for the types it names. Defaults: {"skill": 0.60, "knowledge": 0.50, "task": 0.50} |
None
|
levels
|
bool
|
Accepted for compatibility; currently has no effect. |
False
|
batch_size
|
int
|
Accepted for compatibility; currently has no effect. |
DEFAULT_BATCH_SIZE
|
warnings
|
bool
|
Whether to show warnings |
False
|
concepts
|
list
|
Types to extract: "skills", "knowledge", "tasks", or ["all"]. Defaults to ["skills"]. |
None
|
extract
|
list
|
Backward-compatible alias for |
None
|
return_edges
|
bool
|
If True, return {"nodes": pd.DataFrame, "edges": pd.DataFrame} where "edges" contains ENABLES edges (Knowledge → Task per skill). If False (default), return a plain pd.DataFrame. |
False
|
timing
|
bool
|
Accepted for compatibility with benchmark callers; currently has no effect. |
False
|
output_csv_path
|
str
|
If provided, write normalized results to this CSV path. |
None
|
Returns:
| Type | Description |
|---|---|
pd.DataFrame (when return_edges=False)
|
Normalized mixed-concept rows with: Research ID, Type, Raw Concept, Taxonomy Concept, Taxonomy Description, Taxonomy Source, Source Url, Correlation Coefficient |
dict (when return_edges=True): {"nodes": pd.DataFrame, "edges": pd.DataFrame}
|
|
extract_and_align ¶
extract_and_align(data: DataFrame, id_column: str = 'Research ID', text_columns: List[str] = None, input_type: str = 'job_desc', top_k: Optional[int] = None, similarity_threshold: Optional[float] = None, levels: bool = False, batch_size: int = DEFAULT_BATCH_SIZE, warnings: bool = False, allowed_sources: Optional[List[str]] = None, extract: List[str] = None, return_edges: bool = False, similarity_thresholds: Optional[Dict[str, float]] = None, timing: bool = False, output_csv_path: Optional[str] = None)
Backward-compatible skills-only extraction wrapper.
This method is retained to avoid breaking existing integrations that
still call extract_and_align(...). It always runs the skills path.
Use extract_concepts(...) for mixed concept extraction.