What it does
Credit risk data cleaning and variable screening pipeline for pre-loan modeling
Skills ZICQ category:Agent Workflows datanalysis-credit-risk
Credit risk data cleaning and variable screening pipeline for pre-loan modeling. Use when working with raw credit data that needs quality assessment, missing value analysis, or variable selection before modeling. it covers data loading and formatting, abnormal period filtering, missing rate calculation, high-missing variable removal,low-IV variable filtering, high-PSI variable removal, Null Importance denoising, high-correlation variable removal, and cleaning report generation. Applicable scenarios arecredit risk data cleaning, variable screening, pre-loan modeling preprocessing.
Official URL:skills.sh
Intro in this page language first. The official description stays in its original wording; we do not rewrite SKILL.md.
Credit risk data cleaning and variable screening pipeline for pre-loan modeling
working with raw credit data that needs quality assessment, missing value analysis, or variable selection before modeling. it covers data loading and formatting, abnormal period filtering, missing rate calculation, high-missing variable removal,low-IV variable filtering, high-PS
Per Agent Skills progressive disclosure: name and description load at startup (~100 tokens); the full SKILL.md body loads when the skill activates; scripts/, references/, and assets/ load only as needed. This file's sections: Data Cleaning and Variable Screening; Quick Start; Run the complete data cleaning pipeline; Complete Process Description; Core Functions; Parameter Description.
File analysis: besides SKILL.md, the body references scripts/example.py. Those resources load on demand.
Data Cleaning and Variable ScreeningQuick StartRun the complete data cleaning pipelineComplete Process DescriptionCore FunctionsParameter DescriptionData Loading ParametersOOS Organization ConfigurationAbnormal Month Filtering ParametersMissing Rate ParametersIV ParametersPSI Parameters
Source category:skills.sh agent-skill
namedatanalysis-credit-riskdescriptionscripts/example.pyThese paths are extracted from the text. Check the upstream package to verify the files exist.
Invocation syntax and available tools depend on your Agent client. Client integration guide ↗
Choose the target agent and installation scope, keep referenced package files, then verify the skill appears in the client's catalog.
This skill references supporting files. Retrieve the complete directory from the source; copying SKILL.md alone may leave missing dependencies.
Copy these instructions to a compatible agent and confirm the target directory matches your client.
Install the agent skill "datanalysis-credit-risk" into my project. The full SKILL.md and official description are at https://zicq.com/en/skills/skl-4e9b465c4c410849-Datanalysis-Credit-Risk.html Save it as .cursor/skills/datanalysis-credit-risk/SKILL.md or .claude/skills/datanalysis-credit-risk/SKILL.md and keep the frontmatter name and description exactly as-is. This skill also ships scripts/, references/, or assets/ — fetch the whole folder from https://github.com/github/awesome-copilot instead of creating only a SKILL.md.
Requires Node.js and npx. First inspect the repository's skill list to confirm the name.
npx skills add 'https://github.com/github/awesome-copilot' --list
npx skills add 'https://github.com/github/awesome-copilot' --skill 'datanalysis-credit-risk'
The CLI lets you choose the agent interactively. The default scope is the project; use -g for user scope. Confirm package availability with the discovery command, then use npx skills list to inspect installed skills.
# Run the complete data cleaning pipeline
python ".github/skills/datanalysis-credit-risk/scripts/example.py"
The data cleaning pipeline consists of the following 11 steps, each executed independently without deleting the original data:
| Function | Purpose | Module |
|------|------|----------|
| get_dataset() | Load and format data | references.func |
| org_analysis() | Organization sample analysis | references.func |
| missing_check() | Calculate missing rate | references.func |
| drop_abnormal_ym() | Filter abnormal months | references.analysis |
| drop_highmiss_features() | Drop high missing rate features | references.analysis |
| drop_lowiv_features() | Drop low IV features | references.analysis |
| drop_highpsi_features() | Drop high PSI features | references.analysis |
| drop_highnoise_features() | Null Importance denoising | references.analysis |
| drop_highcorr_features() | Drop high correlation features | references.analysis |
| iv_distribution_by_org() | IV distribution statistics | references.analysis |
| psi_distribution_by_org() | PSI distribution statistics | references.analysis |
| value_ratio_distribution_by_org() | Value ratio distribution statistics | references.analysis |
| export_cleaning_report() | Export cleaning report | references.analysis |
DATA_PATH: Data file path (best are parquet format)DATE_COL: Date column nameY_COL: Label column nameORG_COL: Organization column nameKEY_COLS: Primary key column name listOOS_ORGS: Out-of-sample organization listmin_ym_bad_sample: Minimum bad sample count per month (default 10)min_ym_sample: Minimum total sample count per month (default 500)missing_ratio: Overall missing rate threshold (default 0.6)overall_iv_threshold: Overall IV threshold (default 0.1)org_iv_threshold: Single organization IV threshold (default 0.1)max_org_threshold: Maximum tolerated low IV organization count (default 2)psi_threshold: PSI threshold (default 0.1)max_months_ratio: Maximum unstable month ratio (default 1/3)max_orgs: Maximum unstable organization count (default 6)n_estimators: Number of trees (default 100)max_depth: Maximum tree depth (default 5)gain_threshold: Gain difference threshold (default 50)max_corr: Correlation threshold (default 0.9)top_n_keep: Keep top N features by original gain ranking (default 20)The generated Excel report contains the following sheets:
Agent Workflows
Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that exte…
Agent Workflows
Use the ClawdHub CLI to search, install, update, and publish agent skills from clawdhub.com. Use when you need to fetch new skills on the fl…
Agent Workflows
Orchestrate multi-agent teams with defined roles, task lifecycles, handoff protocols, and review workflows. Use when: (1) Setting up a team …
Agent Workflows
Spec-first, TDD, subagent-driven software development workflow. Use when: (1) building any new feature or app — triggers brainstorm → plan →…