网页元数据和联系信息提取器 API

API ID 13498

将任何 URL 转换为结构化数据:SEO/OpenGraph 元数据 联系人 社交档案 40 多个技术栈签名 SEO 和安全审计 以及 AI 准备好的 Markdown 只需一次快速的 API 调用

303 ms 平均响应

API 文档

端点

请求
返回一个 URL 的完整元数据负载一次性调用 SEO/OpenGraph 元数据 页面健康指标 公共联系人 社交资料 检测到的技术 14 点 SEO 审核 分类的内部/外部链接 Schema.org 结构化数据 RSS/Atom 提要 HTTP 安全头分数 干净的 AI 准备好 Markdown 包含字数和阅读时间 使用可选字段参数仅运行实际需要的提取器 费用与调用匹配的专门端点相同
Endpoint ID: 29644
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29644/extract+metadata
输入参数

提取元数据 — 端点功能

对象 描述
url 必需 The target website URL to analyze (e.g. https://example.com)
fields 可选 Optional comma-separated list of keys to filter the response (e.g. metadata,contacts) — only runs the extractors those fields need
user_agent 可选 Optional custom User-Agent header for the outbound fetch

剩余免费测试请求:3 / 3。


输入参数

url
fields
user_agent
API 示例响应
JSON
{
    "url": "https://example.com",
    "final_url": "https://example.com",
    "status_code": 200,
    "execution_time_ms": 113.63,
    "bot_protection_detected": false,
    "metadata": {
        "title": "Example Domain",
        "description": null,
        "og_image": null,
        "og_type": null,
        "og_url": null,
        "og_video": null,
        "og_locale_alternate": [],
        "keywords": null,
        "author": null,
        "site_name": null,
        "language": "en",
        "favicon": "data:,",
        "favicon_high_res": "https://www.google.com/s2/favicons?domain=example.com&sz=128",
        "canonical_url": null,
        "theme_color": null,
        "robots": null,
        "hreflang_tags": [],
        "h1_tags": [
            "Example Domain"
        ],
        "h1_count": 1,
        "images_count": 0,
        "images_missing_alt_count": 0,
        "links_count": 1,
        "video_embed_code": null,
        "ssl_status": {
            "enabled": true,
            "hsts_active": false,
            "protocol": "HTTPS"
        },
        "viewport": "width=device-width, initial-scale=1",
        "twitter_card": null,
        "summary_snippet": "Example Domain This domain is for use in documentation examples without needing permission. Avoid use in operations.",
        "top_keywords": [
            "domain",
            "use",
            "example",
            "documentation",
            "examples"
        ],
        "content_length_bytes": 559
    },
    "social_links": {
        "instagram": null,
        "github": null,
        "medium": null,
        "threads": null,
        "pinterest": null,
        "facebook": null,
        "twitter": null,
        "gitlab": null,
        "tiktok": null,
        "bluesky": null,
        "mastodon": null,
        "snapchat": null,
        "telegram": null,
        "discord": null,
        "youtube": null,
        "behance": null,
        "reddit": null,
        "vimeo": null,
        "dribbble": null,
        "linkedin": null
    },
    "contacts": {
        "emails": [],
        "phones": []
    },
    "phone_details": [],
    "detected_technologies": [],
    "technology_details": [],
    "rss_feeds": [],
    "json_ld_schemas": [],
    "product_data": null,
    "product_field_confidence": [],
    "quality": {
        "score": 0.9,
        "rendered": false,
        "sources_used": [
            "meta"
        ],
        "warnings": [
            {
                "field": "metadata.description",
                "type": "MISSING_FIELD"
            }
        ]
    },
    "security_score_percentage": 0,
    "seo_score_percentage": 58.3,
    "seo_passed_checks": [
        "Title tag present with optimal length (10-70 chars)",
        "Primary H1 heading present (1 found)",
        "Favicon icon present",
        "HTTPS secure protocol active",
        "HTML lang attribute present ('en')",
        "Responsive viewport meta tag present",
        "Page is indexable (no noindex directive)"
    ],
    "seo_warnings": [
        "Missing <meta name='description'> tag",
        "Missing <link rel='canonical'> tag",
        "Missing og:image social preview tag",
        "Missing Twitter Card meta tag",
        "No structured data (JSON-LD) found"
    ],
    "seo_checks": [
        {
            "check": "title",
            "passed": true,
            "severity": "critical",
            "evidence": "Title tag present with optimal length (10-70 chars)"
        },
        {
            "check": "meta_description",
            "passed": false,
            "severity": "critical",
            "evidence": "Missing <meta name='description'> tag"
        },
        {
            "check": "canonical",
            "passed": false,
            "severity": "important",
            "evidence": "Missing <link rel='canonical'> tag"
        },
        {
            "check": "h1_present",
            "passed": true,
            "severity": "critical",
            "evidence": "Primary H1 heading present (1 found)"
        },
        {
            "check": "og_image",
            "passed": false,
            "severity": "important",
            "evidence": "Missing og:image social preview tag"
        },
        {
            "check": "favicon",
            "passed": true,
            "severity": "minor",
            "evidence": "Favicon icon present"
        },
        {
            "check": "https",
            "passed": true,
            "severity": "critical",
            "evidence": "HTTPS secure protocol active"
        },
        {
            "check": "lang_attribute",
            "passed": true,
            "severity": "important",
            "evidence": "HTML lang attribute present ('en')"
        },
        {
            "check": "viewport",
            "passed": true,
            "severity": "important",
            "evidence": "Responsive viewport meta tag present"
        },
        {
            "check": "indexability",
            "passed": true,
            "severity": "critical",
            "evidence": "Page is indexable (no noindex directive)"
        },
        {
            "check": "twitter_card",
            "passed": false,
            "severity": "minor",
            "evidence": "Missing Twitter Card meta tag"
        },
        {
            "check": "structured_data",
            "passed": false,
            "severity": "minor",
            "evidence": "No structured data (JSON-LD) found"
        }
    ],
    "internal_links": [],
    "external_links": [
        "https://iana.org/domains/example"
    ],
    "total_internal_count": 0,
    "total_external_count": 1,
    "word_count": 20,
    "reading_time_minutes": 0.1,
    "markdown_content": "# Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n\nLearn more"
}
提取元数据 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29644/extract+metadata?url=https://example.com' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
轻量级终端优化用于链接预览卡片(社交卡片/展开卡片)返回给定网址的标题 描述 og_image 网站图标 高清网站图标 网站名称 和语言 适合聊天应用程序 书签工具 和分享预览
Endpoint ID: 29645
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29645/extract+link+preview
输入参数

提取链接预览 — 端点功能

对象 描述
url 必需 The target URL to generate a link preview for

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://example.com",
    "final_url": "https://example.com",
    "status_code": 200,
    "execution_time_ms": 19.13,
    "bot_protection_detected": false,
    "title": "Example Domain",
    "description": null,
    "og_image": null,
    "favicon": "data:,",
    "favicon_high_res": "https://www.google.com/s2/favicons?domain=example.com&sz=128",
    "site_name": null,
    "language": "en"
}
提取链接预览 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29645/extract+link+preview?url=https://example.com' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
用于发现公共联系信号的专用端点 返回在页面上找到的公共电子邮件 电话号码和官方社交媒体资料网址 仅呈现原始公共信号 不包含有关公司或个人的验证数据 也不识别它们属于谁
Endpoint ID: 29646
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29646/extraer+contactos
输入参数

提取联系人 — 端点功能

对象 描述
url 必需 The target URL to extract contact information and social handles from

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://www.eff.org/about/contact",
    "final_url": "https://www.eff.org/about/contact",
    "status_code": 200,
    "execution_time_ms": 93.62,
    "bot_protection_detected": false,
    "emails": [
        "[email protected]",
        "[email protected]",
        "[email protected]",
        "[email protected]",
        "[email protected]",
        "[email protected]"
    ],
    "phones": [],
    "phone_details": [],
    "social_links": {
        "reddit": null,
        "gitlab": null,
        "medium": null,
        "twitter": null,
        "youtube": "https://www.youtube.com/efforg",
        "vimeo": null,
        "telegram": null,
        "linkedin": "https://www.linkedin.com/company/EFF",
        "bluesky": "https://bsky.app/profile/eff.org",
        "behance": null,
        "mastodon": "https://mastodon.social/@eff",
        "threads": "https://www.threads.net/@efforg",
        "tiktok": "https://www.tiktok.com/@efforg",
        "instagram": "https://www.instagram.com/efforg/",
        "facebook": "https://www.facebook.com/eff",
        "github": null,
        "discord": null,
        "dribbble": null,
        "pinterest": null,
        "snapchat": null
    }
}
提取联系人 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29646/extraer+contactos?url=https://www.eff.org/about/contact' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
专用端点用于技术智能和CMS审计 检测40多个框架和CMS签名(WordPress Shopify React Next.js Vue Angular Stripe GA4等)具有结构上下文感知匹配和每个匹配的置信度得分
Endpoint ID: 29647
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29647/extract+tech+stack
输入参数

提取技术栈 — 端点功能

对象 描述
url 必需 The target URL to inspect for CMS and technology stack signatures

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://techcrunch.com",
    "final_url": "https://techcrunch.com",
    "status_code": 200,
    "execution_time_ms": 83.98,
    "bot_protection_detected": false,
    "detected_technologies": [
        "WordPress",
        "Google Tag Manager",
        "Cloudflare"
    ],
    "technology_details": [
        {
            "name": "WordPress",
            "confidence": 0.9,
            "evidence": [
                "wp-content",
                "wp-includes"
            ],
            "category": "cms"
        },
        {
            "name": "Google Tag Manager",
            "confidence": 0.75,
            "evidence": [
                "googletagmanager.com/gtm.js"
            ],
            "category": "analytics"
        },
        {
            "name": "Cloudflare",
            "confidence": 0.75,
            "evidence": [
                "cloudflare.com"
            ],
            "category": "hosting"
        }
    ]
}
提取技术栈 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29647/extract+tech+stack?url=https://techcrunch.com' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
专用端点用于Schema.org和JSON-LD结构化数据提取 返回解析的产品定价 电子商务评论 文章模式 活动细节和组织元数据 跨来源核对价格 货币 可用性和品牌并标记冲突
Endpoint ID: 29648
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29648/extract+schema
输入参数

提取模式 — 端点功能

对象 描述
url 必需 The target URL to extract Schema.org JSON-LD structured data from

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://techcrunch.com",
    "final_url": "https://techcrunch.com",
    "status_code": 200,
    "execution_time_ms": 0.69,
    "bot_protection_detected": false,
    "json_ld_count": 1,
    "json_ld_schemas": [
        {
            "@context": "https://schema.org",
            "@graph": [
                {
                    "@type": "CollectionPage",
                    "@id": "https://techcrunch.com/",
                    "url": "https://techcrunch.com/",
                    "name": "TechCrunch | Startup and Technology News",
                    "isPartOf": {
                        "@id": "https://techcrunch.com/#website"
                    },
                    "about": {
                        "@id": "https://techcrunch.com/#organization"
                    },
                    "description": "TechCrunch | Reporting on the business of technology, startups, venture capital funding, and Silicon Valley",
                    "breadcrumb": {
                        "@id": "https://techcrunch.com/#breadcrumb"
                    },
                    "inLanguage": "en-US"
                },
                {
                    "@type": "BreadcrumbList",
                    "@id": "https://techcrunch.com/#breadcrumb",
                    "itemListElement": [
                        {
                            "@type": "ListItem",
                            "position": 1,
                            "name": "Home"
                        }
                    ]
                },
                {
                    "@type": "WebSite",
                    "@id": "https://techcrunch.com/#website",
                    "url": "https://techcrunch.com/",
                    "name": "TechCrunch",
                    "description": "Startup and Technology News",
                    "publisher": {
                        "@id": "https://techcrunch.com/#organization"
                    },
                    "alternateName": "TC",
                    "potentialAction": [
                        {
                            "@type": "SearchAction",
                            "target": {
                                "@type": "EntryPoint",
                                "urlTemplate": "https://techcrunch.com/?s={search_term_string}"
                            },
                            "query-input": {
                                "@type": "PropertyValueSpecification",
                                "valueRequired": true,
                                "valueName": "search_term_string"
                            }
                        }
                    ],
                    "inLanguage": "en-US"
                },
                {
                    "@type": "Organization",
                    "@id": "https://techcrunch.com/#organization",
                    "name": "TechCrunch",
                    "alternateName": "TC",
                    "url": "https://techcrunch.com/",
                    "logo": {
                        "@type": "ImageObject",
                        "inLanguage": "en-US",
                        "@id": "https://techcrunch.com/#/schema/logo/image/",
                        "url": "https://techcrunch.com/wp-content/uploads/2018/04/tc-logo-2018-square-reverse2x.png?resize=1200,1200",
                        "contentUrl": "https://techcrunch.com/wp-content/uploads/2018/04/tc-logo-2018-square-reverse2x.png?resize=1200,1200",
                        "width": 1200,
                        "height": 1200,
                        "caption": "TechCrunch"
                    },
                    "image": {
                        "@id": "https://techcrunch.com/#/schema/logo/image/"
                    },
                    "sameAs": [
                        "https://www.facebook.com/techcrunch",
                        "https://x.com/TechCrunch",
                        "https://mstdn.social/@TechCrunch",
                        "https://bsky.app/profile/techcrunch.com",
                        "https://www.threads.net/@techcrunch"
                    ]
                }
            ]
        }
    ]
}
提取模式 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29648/extract+schema?url=https://techcrunch.com' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
用于HTTP安全头审计的专用端点检查HSTS 内容安全策略 X-Frame-Options X-Content-Type-Options和引荐政策然后为目标站点计算一个评级百分比安全分数
Endpoint ID: 29649
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29649/extract+security+headers
输入参数

提取安全头部 — 端点功能

对象 描述
url 必需 The target URL to audit HTTP security headers

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://google.com",
    "final_url": "https://www.google.com/",
    "status_code": 200,
    "execution_time_ms": 226.19,
    "bot_protection_detected": false,
    "security_score_percentage": 36.7,
    "security_headers": {
        "strict_transport_security": "max-age=31536000",
        "content_security_policy": null,
        "x_frame_options": "SAMEORIGIN",
        "x_content_type_options": null,
        "referrer_policy": null,
        "permissions_policy": "unload=()"
    },
    "security_header_grades": {
        "strict_transport_security": "reasonable",
        "content_security_policy": "missing",
        "x_frame_options": "strong",
        "x_content_type_options": "missing",
        "referrer_policy": "missing",
        "permissions_policy": "reasonable"
    }
}
提取安全头部 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29649/extract+security+headers?url=https://google.com' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
专用端点用于AI代理、ChatGPT、Claude和RAG管道。去除噪音(广告、导航、页脚、脚本)并将网页文章文本转换为干净、结构化的Markdown,包括字数和阅读时间
Endpoint ID: 29650
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29650/extract+clean+markdown
输入参数

提取干净的Markdown — 端点功能

对象 描述
url 必需 The target article or webpage URL to extract clean LLM-ready Markdown from

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://en.wikipedia.org/wiki/Web_scraping",
    "final_url": "https://en.wikipedia.org/wiki/Web_scraping",
    "status_code": 200,
    "execution_time_ms": 129.45,
    "bot_protection_detected": false,
    "title": "Web scraping - Wikipedia",
    "word_count": 2987,
    "reading_time_minutes": 14.9,
    "summary_snippet": "Jump to content \n \n\t \n\t\t \n\t\t\t \n\t\t\t\t\n \n\t \n\t \n\n Main menu \n\t \n\t \n\n\n\t\t\t\t \n\t\t\n \n\t \n\t Main menu \n\t move to sidebar \n\t hide \n \n\n\t\n \n\t \n\t\tNavigation\n\t \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t Main page Contents Current events Random article About Wikipedia Contact us \n\t\t \n\t\t\n\t \n \n\n\t\n \n\t \n\t\tContribute\n\t \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t Help Learn to edit Community portal Recent changes Upload file Special pages \n\t\t \n\t\t\n\t \n \n\n \n\n\t\t\t\t \n\n\t \n \n\n\t\t \n\t\t\t\n \n\t \n\t \n\t\t \n\t\t \n\t \n \n\n\t\t \n\t\t \n\t\t\t\n \n\t \n\n Search \n\t \n\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t \n\t\t\t\t Search \n\t\t\t \n\t\t \n\t \n \n\n\t\t\t \n\t \n\t\n \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t\n \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t \n\t\t\n \n\t \n\t \n\n Appearance \n\t \n\t \n\n\n\t\t\t \n\t\t\t\t\n\t\t\t \n\t\t\n\t \n \n\n\t \n\t\n \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t\n \n\t \n\t\t\n\t\t \n\t\t\t Donate \n \n Create account \n \n Log in \n \n\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t \n\t\n \n\t \n\t \n\n Personal tools \n\t \n\t \n\n\n\t\t\n \n\t \n\t\t\n\t\t \n\t\t\t \n\n Donate \n \n \n\n Create account \n \n \n\n Log in \n \n\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t\n\t \n \n\n \n\n\t\t \n\t \n \n \n\t \n\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\n\t\t\t\t \n\t\t \n\t\t \n\t \n\t \n\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t \n\t \n\t Contents \n\t move to sidebar \n\t hide \n \n\n\n\t \n\t\t \n\t\t\t \n\t\t\t\t (Top) \n\t\t\t \n\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t 1 \n\t\t\t\t History \n\t\t\t \n\t\t \n\t\t\n\t\t \n\t\t \n\t \n\t \n\t\t \n\t\t\t \n\t\t\t\t 2 \n\t\t\t\t Techniques \n\t\t\t \n\t\t \n\t\t\n\t\t\t \n\t\t\t\t \n\t\t\t\t Toggle Techniques subsection \n\t\t\t \n\t\t\n\t\t \n\t\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.1 \n\t\t\t\t\t Human copy-and-paste \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.2 \n\t\t\t\t\t Text pattern matching \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.3 \n\t\t\t\t\t HTTP programming \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.4 \n\t\t\t\t\t HTML parsing \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.5 \n\t\t\t\t\t DOM parsing \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.6 \n\t\t\t\t\t Vertical aggregation \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.7 \n\t\t\t\t\t Semantic annotation recognizing \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 2.8 \n\t\t\t\t\t Computer vision web-page analysis \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t \n\t \n\t \n\t\t \n\t\t\t \n\t\t\t\t 3 \n\t\t\t\t Legal issues \n\t\t\t \n\t\t \n\t\t\n\t\t\t \n\t\t\t\t \n\t\t\t\t Toggle Legal issues subsection \n\t\t\t \n\t\t\n\t\t \n\t\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 3.1 \n\t\t\t\t\t United States \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 3.2 \n\t\t\t\t\t European Union \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 3.3 \n\t\t\t\t\t Australia \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t 3.4 \n\t\t\t\t\t India \n\t\t\t\t \n\t\t\t \n\t\t\t\n\t\t\t \n\t\t\t \n\t\t \n\t \n\t \n\t \n\t\t \n\t\t\t \n\t\t\t\t 4 \n\t\t\t\t Methods to prevent web scraping \n\t\t\t \n\t\t \n\t\t\n\t\t \n\t\t \n\t \n\t \n\t\t \n\t\t\t \n\t\t\t\t 5 \n\t\t\t\t See also \n\t\t\t \n\t\t \n\t\t\n\t\t \n\t\t \n\t \n\t \n\t\t \n\t\t\t \n\t\t\t\t 6 \n\t\t\t\t References \n\t\t\t \n\t\t \n\t\t\n\t\t \n\t\t \n\t \n \n \n\n\t\t\t\t\t \n\t\t \n\t\t\t \n\t\t \n\t\t \n\t\t\t \n\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t\t\n \n\t \n\t \n\n Toggle the table of contents \n\t \n\t \n\n\n\t\t\t\t\t\t\t \n\t\t\t \n\t\t\n\t \n \n\n\t\t\t\t\t \n\t\t\t\t\t Web scraping \n\t\t\t\t\t\t\t\n \n\t \n\t \n\n 22 languages \n\t \n\t \n\n\t\t \n\t\t\t\n\t\t\t \n\t\t\t\t\n\t\t\t\t العربية الدارجة Català Čeština Deutsch Español Euskara فارسی Français Bahasa Indonesia Íslenska Italiano 日本語 한국어 Latviešu Nederlands Português Русский Türkçe Українська 粵語 中文 \n\t\t\t \n\t\t\t Edit links \n\t\t \n\n\t \n \n \n\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t\t\t \n\t\t\t\t\t\t\t\t\n \n\t \n\t\t\n\t\t \n\t\t\t Article \n \n Talk \n \n\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t\t\t\t\t\t\t\t\n \n\t \n\t English \n\t \n\t \n\n\n\t\t\t\t\t\n \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t\t\t\t\n\t \n \n\n\t\t\t\t\t\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t\t\t \n\t\t\t    \n \n\t \n\t\t\n\t\t \n\t\t\t Read \n \n Edit \n \n View history \n \n\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n\t\t\t \n\n\t\t\t\t\t\t\t \n\t\t\t\t\t\t\t\t\n \n\t \n\t \n\n Tools \n\t \n\t \n\n\n\t\t\t\t\t\t\t\t\t \n\t\t\t\t\t\t\n \n\t \n\t Tools \n\t move to sidebar \n\t hide \n \n\n\t\n \n\t \n\t\tActions\n\t \n\t \n\t\t\n\t\t \n\t\t\t \n\n Read \n \n \n\n Edit \n \n \n\n View history \n \n\n\t\t\t\n\t\t \n\t\t\n\t \n \n\n \n\t \n\t\tGeneral\n\t \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t What links here Related changes Upload file Permanent link Page information Cite this page Get shortened URL Switch to legacy parser \n\t\t \n\t\t\n\t \n \n\n \n\t \n\t\tPrint/export\n\t \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t Download as PDF Printable version \n\t\t \n\t\t\n\t \n \n\n \n\t \n\t\tIn other projects\n\t \n\t \n\t\t\n\t\t \n\t\t\t\n\t\t\t Wikimedia Commons Wikidata item \n\t\t \n\t\t\n\t \n \n\n \n\n\t\t\t\t\t\t\t\t\t \n\t\t\t\t\n\t \n \n\n\t\t\t\t\t\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t \n\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t\t\t \n\t\t\t\t\n\t\t\t\t\t\t\t \n\t\t \n\t\t\t\t\t\t \n\t\t\t\t\t\t\t \n\t\t\t\t \n\t \n\t Appearance \n\t move to sidebar \n\t hide \n \n\n\n \n\n\t\t\t\t\t\t\t \n\t\t \n\t\t\t\t\t \n\t\t\t\t \n\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t\t\t \n\t\t \n\n\t\t\t\t\t\t From Wikipedia, the free encyclopedia \n\t\t\t\t\t \n\t\t\t\t\t \n\t\t\t\t\t\n\t\t\t\t\t\n\t\t\t\t\t \n Method of extracting data from websites \n For broader coverage of this topic, see  Data scraping . \"Web scraper\" redirects here.",
    "top_keywords": [
        "web",
        "scraping",
        "data",
        "pages",
        "edit"
    ],
    "_note": "Response truncated for documentation purposes"
}
提取干净的Markdown — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29650/extract+clean+markdown?url=https://en.wikipedia.org/wiki/Web_scraping' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
用于自动化14点SEO诊断审计的专用端点 评估标题标签 元描述 规范网址 H1标题 OpenGraph图像 网站图标 图像ALT覆盖 HTTPS lang属性 视口 可索引性 多个H1 Twitter卡片 和结构化数据 返回一个0-100%的评分结果并附带可操作的警告
Endpoint ID: 29651
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29651/extract+seo+audit
输入参数

提取SEO审核 — 端点功能

对象 描述
url 必需 The target URL to perform an automated 14-point SEO diagnostic audit on

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://example.com",
    "final_url": "https://example.com",
    "status_code": 200,
    "execution_time_ms": 32.04,
    "bot_protection_detected": false,
    "seo_score_percentage": 58.3,
    "passed_checks": [
        "Title tag present with optimal length (10-70 chars)",
        "Primary H1 heading present (1 found)",
        "Favicon icon present",
        "HTTPS secure protocol active",
        "HTML lang attribute present ('en')",
        "Responsive viewport meta tag present",
        "Page is indexable (no noindex directive)"
    ],
    "warnings": [
        "Missing <meta name='description'> tag",
        "Missing <link rel='canonical'> tag",
        "Missing og:image social preview tag",
        "Missing Twitter Card meta tag",
        "No structured data (JSON-LD) found"
    ],
    "checks": [
        {
            "check": "title",
            "passed": true,
            "severity": "critical",
            "evidence": "Title tag present with optimal length (10-70 chars)"
        },
        {
            "check": "meta_description",
            "passed": false,
            "severity": "critical",
            "evidence": "Missing <meta name='description'> tag"
        },
        {
            "check": "canonical",
            "passed": false,
            "severity": "important",
            "evidence": "Missing <link rel='canonical'> tag"
        },
        {
            "check": "h1_present",
            "passed": true,
            "severity": "critical",
            "evidence": "Primary H1 heading present (1 found)"
        },
        {
            "check": "og_image",
            "passed": false,
            "severity": "important",
            "evidence": "Missing og:image social preview tag"
        },
        {
            "check": "favicon",
            "passed": true,
            "severity": "minor",
            "evidence": "Favicon icon present"
        },
        {
            "check": "https",
            "passed": true,
            "severity": "critical",
            "evidence": "HTTPS secure protocol active"
        },
        {
            "check": "lang_attribute",
            "passed": true,
            "severity": "important",
            "evidence": "HTML lang attribute present ('en')"
        },
        {
            "check": "viewport",
            "passed": true,
            "severity": "important",
            "evidence": "Responsive viewport meta tag present"
        },
        {
            "check": "indexability",
            "passed": true,
            "severity": "critical",
            "evidence": "Page is indexable (no noindex directive)"
        },
        {
            "check": "twitter_card",
            "passed": false,
            "severity": "minor",
            "evidence": "Missing Twitter Card meta tag"
        },
        {
            "check": "structured_data",
            "passed": false,
            "severity": "minor",
            "evidence": "No structured data (JSON-LD) found"
        }
    ]
}
提取SEO审核 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29651/extract+seo+audit?url=https://example.com' --header 'Authorization: Bearer YOUR_API_KEY' 


    
请求
专用端点用于超链接提取和域名分类 将页面上的所有链接分类为内部链接(同一域)和外部链接(第三方网站) 每页最多100个链接 以及RSS/Atom订阅源发现
Endpoint ID: 29652
GET https://docs.zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29652/extract+links
输入参数

提取链接 — 端点功能

对象 描述
url 必需 The target URL to extract and classify internal vs external hyperlinks from

剩余免费测试请求:3 / 3。


输入参数

url
API 示例响应
JSON
{
    "url": "https://en.wikipedia.org/wiki/Web_scraping",
    "final_url": "https://en.wikipedia.org/wiki/Web_scraping",
    "status_code": 200,
    "execution_time_ms": 51.72,
    "bot_protection_detected": false,
    "total_links_count": 85,
    "internal_links_count": 50,
    "external_links_count": 35,
    "internal_links": [
        "https://en.wikipedia.org/wiki/Main_Page",
        "https://en.wikipedia.org/wiki/Wikipedia:Contents",
        "https://en.wikipedia.org/wiki/Portal:Current_events",
        "https://en.wikipedia.org/wiki/Special:Random",
        "https://en.wikipedia.org/wiki/Wikipedia:About",
        "https://en.wikipedia.org/wiki/Wikipedia:Contact_us",
        "https://en.wikipedia.org/wiki/Help:Contents",
        "https://en.wikipedia.org/wiki/Help:Introduction",
        "https://en.wikipedia.org/wiki/Wikipedia:Community_portal",
        "https://en.wikipedia.org/wiki/Special:RecentChanges",
        "https://en.wikipedia.org/wiki/Wikipedia:File_upload_wizard",
        "https://en.wikipedia.org/wiki/Special:SpecialPages",
        "https://en.wikipedia.org/wiki/Special:Search",
        "https://en.wikipedia.org/w/index.php?title=Special:CreateAccount&returnto=Web+scraping",
        "https://en.wikipedia.org/w/index.php?title=Special:UserLogin&returnto=Web+scraping",
        "https://en.wikipedia.org/wiki/Web_scraping",
        "https://en.wikipedia.org/wiki/Talk:Web_scraping",
        "https://en.wikipedia.org/w/index.php?title=Web_scraping&action=edit",
        "https://en.wikipedia.org/w/index.php?title=Web_scraping&action=history",
        "https://en.wikipedia.org/wiki/Special:WhatLinksHere/Web_scraping",
        "https://en.wikipedia.org/wiki/Special:RecentChangesLinked/Web_scraping",
        "https://en.wikipedia.org/wiki/Wikipedia:File_Upload_Wizard",
        "https://en.wikipedia.org/w/index.php?title=Web_scraping&oldid=1368486598",
        "https://en.wikipedia.org/w/index.php?title=Web_scraping&action=info",
        "https://en.wikipedia.org/w/index.php?title=Special:CiteThisPage&page=Web_scraping&id=1368486598&wpFormIdentifier=titleform",
        "https://en.wikipedia.org/w/index.php?title=Special:UrlShortener&url=https%3A%2F%2Fen.wikipedia.org%2Fwiki%2FWeb_scraping",
        "https://en.wikipedia.org/w/index.php?title=Web_scraping&useparsoid=0",
        "https://en.wikipedia.org/w/index.php?title=Special:DownloadAsPdf&page=Web_scraping&action=show-download-screen",
        "https://en.wikipedia.org/w/index.php?title=Web_scraping&printable=yes",
        "https://en.wikipedia.org/wiki/Data_scraping",
        "https://en.wikipedia.org/wiki/Scraper_site",
        "https://en.wikipedia.org/wiki/File:Question_book-new.svg",
        "https://en.wikipedia.org/wiki/Wikipedia:Verifiability",
        "https://en.wikipedia.org/wiki/Special:EditPage/Web_scraping",
        "https://en.wikipedia.org/wiki/Help:Referencing_for_beginners",
        "https://en.wikipedia.org/wiki/Wikipedia:Verifiability#Burden_of_evidence",
        "https://en.wikipedia.org/wiki/Help:Maintenance_template_removal",
        "https://en.wikipedia.org/wiki/Data_extraction",
        "https://en.wikipedia.org/wiki/Website",
        "https://en.wikipedia.org/wiki/World_Wide_Web",
        "https://en.wikipedia.org/wiki/Hypertext_Transfer_Protocol",
        "https://en.wikipedia.org/wiki/Internet_bot",
        "https://en.wikipedia.org/wiki/Web_crawler",
        "https://en.wikipedia.org/wiki/Database",
        "https://en.wikipedia.org/wiki/Spreadsheet",
        "https://en.wikipedia.org/wiki/Data_retrieval",
        "https://en.wikipedia.org/wiki/Data_analysis",
        "https://en.wikipedia.org/wiki/Parsing",
        "https://en.wikipedia.org/wiki/Contact_scraping",
        "https://en.wikipedia.org/wiki/Web_indexing"
    ],
    "external_links": [
        "https://donate.wikimedia.org/?wmf_source=donate&wmf_medium=sidebar&wmf_campaign=en.wikipedia.org&uselang=en",
        "https://ar.wikipedia.org/wiki/%D8%AA%D8%AC%D8%B1%D9%8A%D9%81_%D9%88%D9%8A%D8%A8",
        "https://ary.wikipedia.org/wiki/%D8%AA%D8%BA%D8%B1%D8%A7%D9%81_%D9%84%D9%88%D9%8A%D8%A8",
        "https://ca.wikipedia.org/wiki/Web_scraping",
        "https://cs.wikipedia.org/wiki/Web_scraping",
        "https://de.wikipedia.org/wiki/Screen_Scraping",
        "https://es.wikipedia.org/wiki/Web_scraping",
        "https://eu.wikipedia.org/wiki/Web_scraping",
        "https://fa.wikipedia.org/wiki/%D8%AA%D8%B1%D8%A7%D8%B4%DB%8C%D8%AF%D9%86_%D9%88%D8%A8",
        "https://fr.wikipedia.org/wiki/Web_scraping",
        "https://id.wikipedia.org/wiki/Penggalian_web",
        "https://is.wikipedia.org/wiki/Vefs%C3%B6fnun",
        "https://it.wikipedia.org/wiki/Web_scraping",
        "https://ja.wikipedia.org/wiki/%E3%82%A6%E3%82%A7%E3%83%96%E3%82%B9%E3%82%AF%E3%83%AC%E3%82%A4%E3%83%94%E3%83%B3%E3%82%B0",
        "https://ko.wikipedia.org/wiki/%EC%9B%B9_%EC%8A%A4%ED%81%AC%EB%9E%98%ED%95%91",
        "https://lv.wikipedia.org/wiki/Rasmo%C5%A1ana",
        "https://nl.wikipedia.org/wiki/Scrapen",
        "https://pt.wikipedia.org/wiki/Web_scraping",
        "https://ru.wikipedia.org/wiki/%D0%92%D0%B5%D0%B1-%D1%81%D0%BA%D1%80%D0%B5%D0%B9%D0%BF%D0%B8%D0%BD%D0%B3",
        "https://tr.wikipedia.org/wiki/Web_kaz%C4%B1ma",
        "https://uk.wikipedia.org/wiki/Web_scraping",
        "https://zh-yue.wikipedia.org/wiki/%E7%B6%B2%E9%A0%81%E5%88%AE%E6%96%99",
        "https://zh.wikipedia.org/wiki/%E7%BD%91%E9%A1%B5%E6%8A%93%E5%8F%96",
        "https://www.wikidata.org/wiki/Special:EntityPage/Q665452#sitelinks-wikipedia",
        "https://commons.wikimedia.org/wiki/Category:Web_scraping",
        "https://www.wikidata.org/wiki/Special:EntityPage/Q665452",
        "https://www.google.com/search?as_eq=wikipedia&q=%22Web+scraping%22",
        "https://www.google.com/search?tbm=nws&q=%22Web+scraping%22+-wikipedia&tbs=ar:1",
        "https://www.google.com/search?&q=%22Web+scraping%22&tbs=bkt:s&tbm=bks",
        "https://www.google.com/search?tbs=bks:1&q=%22Web+scraping%22+-wikipedia",
        "https://scholar.google.com/scholar?q=%22Web+scraping%22",
        "https://www.jstor.org/action/doBasicSearch?Query=%22Web+scraping%22&acc=on&wc=on",
        "https://en.wikiversity.org/wiki/",
        "https://en.wikibooks.org/wiki/",
        "https://en.wikivoyage.org/wiki/"
    ]
}
提取链接 — 代码片段

curl --location --request GET 'https://zylalabs.com/api/13498/web+metadata+and+contact+extractor+api/29652/extract+links?url=https://en.wikipedia.org/wiki/Web_scraping' --header 'Authorization: Bearer YOUR_API_KEY' 


    

API 访问密钥和身份验证

注册后,每个开发者都会被分配一个个人 API 访问密钥,这是一个唯一的字母和数字组合,用于访问我们的 API 端点。要使用 网页元数据和联系信息提取器 API 进行身份验证,只需在 Authorization 标头中包含您的 bearer token。

标头
标头 描述
授权 必需 应为 Bearer access_key. 订阅后,请查看上方的"您的 API 访问密钥"。

无长期承诺。随时升级、降级或取消。 免费试用包括最多 50 个请求。

🚀 企业版套餐

起价
$ 10,000/年


  • 自定义数量
  • 自定义速率限制
  • 专业客户支持
  • 实时 API 监控

概览

将任何 URL 转换为结构化数据:SEO/OpenGraph 元数据、联系人、社交资料、40 多个技术栈签名、SEO 和安全审核,以及适合 AI 的 Markdown — 只需一次快速的 API 调用

网页元数据和联系信息提取器 API FAQs

每个端点返回特定的结构化数据 例如 提取元数据端点提供SEO元数据 健康指标和社交资料 而提取联系人端点返回公共电子邮件和社交媒体链接 其他端点则专注于技术栈检测 安全审计和清理Markdown提取

关键字段因端点而异。对于提取元数据,字段包括标题、描述和语言。提取联系人返回电子邮件和社交链接。提取SEO审核提供分数和警告。每个端点的响应都根据其特定功能量身定制

响应数据采用JSON格式结构,顶层对象包含有关请求的元数据(如URL和状态代码),嵌套对象用于特定数据。例如,提取元数据包括一个"metadata"对象,字段如标题和语言

参数因终端而异 对于提取元数据,您可以使用可选的“字段”参数来指定要提取的数据,从而优化响应 其他终端可能具有针对其特定数据提取需求的唯一参数

典型的使用案例包括SEO分析 内容创作和技术栈审核 例如营销人员可以使用提取SEO审核进行网站优化 而开发人员则可以利用提取技术栈识别网站上使用的框架

数据准确性通过结构化提取方法和验证检查得以保持 每个端点使用特定算法来识别和验证数据 例如检测技术或审核安全头部 确保可靠的输出

用户可以将返回的数据集成到用于分析、报告或内容生成的应用程序中。例如,来自提取干净Markdown端点的干净Markdown可以直接用于人工智能应用程序或内容管理系统

标准数据模式包括具有一致字段名称的结构化JSON响应,适用于相似的端点。例如,大多数端点返回"status_code"和"execution_time_ms",让用户可以评估响应的成功与性能

一般常见问题

Zyla API Hub 就像一个大型 API 商店,您可以在一个地方找到数千个 API。我们还为所有 API 提供专门支持和实时监控。注册后,您可以选择要使用的 API。请记住,每个 API 都需要自己的订阅。但如果您订阅多个 API,您将为所有这些 API 使用相同的密钥,使事情变得更简单。
价格以 USD(美元)、EUR(欧元)、CAD(加元)、AUD(澳元)和 GBP(英镑)列出。我们接受所有主要的借记卡和信用卡。我们的支付系统使用最新的安全技术,由 Stripe 提供支持,Stripe 是世界上最可靠的支付公司之一。如果您在使用卡片付款时遇到任何问题,请通过 [email protected]

此外,如果您已经以这些货币中的任何一种(USD、EUR、CAD、AUD、GBP)拥有有效订阅,该货币将保留用于后续订阅。只要您没有任何有效订阅,您可以随时更改货币。
定价页面上显示的本地货币基于您 IP 地址的国家/地区,仅供参考。实际价格以 USD(美元)为单位。当您付款时,即使您在我们的网站上看到以本地货币显示的等值金额,您的卡片对账单上也会以美元显示费用。这意味着您不能直接使用本地货币付款。
有时,银行可能会因其欺诈保护设置而拒绝收费。我们建议您首先联系您的银行,检查他们是否阻止了我们的收费。此外,您可以访问账单门户并更改关联的卡片以进行付款。如果这些方法不起作用并且您需要进一步帮助,请通过 [email protected]
价格由月度或年度订阅决定,具体取决于所选计划。
API 调用根据成功请求从您的计划中扣除。每个计划都包含您每月可以进行的特定数量的调用。只有成功的调用(由状态 200 响应指示)才会计入您的总数。这确保失败或不完整的请求不会影响您的月度配额。
Zyla API Hub 采用月度订阅系统。您的计费周期将从您购买付费计划的那一天开始,并在下个月的同一日期续订。因此,如果您想避免未来的费用,请提前取消订阅。
要升级您当前的订阅计划,只需转到 API 的定价页面并选择您要升级到的计划。升级将立即生效,让您立即享受新计划的功能。请注意,您之前计划中的任何剩余调用都不会转移到新计划,因此在升级时请注意这一点。您将被收取新计划的全部金额。
要检查您本月剩余多少 API 调用,请参考响应标头中的 "X-Zyla-API-Calls-Monthly-Remaining" 字段。例如,如果您的计划允许每月 1,000 个请求,而您已使用 100 个,则响应标头中的此字段将显示 900 个剩余调用。
要查看您的计划允许的最大 API 请求数,请检查 "X-Zyla-RateLimit-Limit" 响应标头。例如,如果您的计划包括每月 1,000 个请求,此标头将显示 1,000。
"X-Zyla-RateLimit-Reset" 标头显示您的速率限制重置之前的秒数。这告诉您何时您的请求计数将重新开始。例如,如果它显示 3,600,则意味着还有 3,600 秒直到限制重置。
是的,您可以随时通过访问您的账户并在账单页面上选择取消选项来取消您的计划。请注意,升级、降级和取消会立即生效。此外,取消后,您将不再有权访问该服务,即使您的配额中还有剩余调用。
为了让您有机会在没有任何承诺的情况下体验我们的 API,我们提供 7 天免费试用,允许您免费进行最多 50 次 API 调用。此试用只能使用一次,因此我们建议将其应用于您最感兴趣的 API。虽然我们的大多数 API 都提供免费试用,但有些可能不提供。试用在 7 天后或您进行了 50 次请求后结束,以先发生者为准。如果您在试用期间达到 50 次请求限制,您需要"开始您的付费计划"以继续发出请求。您可以在个人资料中的订阅 -> 选择您订阅的 API -> 定价标签下找到"开始您的付费计划"按钮。或者,如果您在第 7 天之前不取消订阅,您的免费试用将结束,您的计划将自动计费,授予您访问计划中指定的所有 API 调用的权限。请记住这一点以避免不必要的费用。
7 天后,您将被收取试用期间订阅的计划的全额费用。因此,在试用期结束前取消很重要。因忘记及时取消而提出的退款请求不被接受。
当您订阅 API 免费试用时,您可以进行最多 50 次 API 调用。如果您希望超出此限制进行额外的 API 调用,API 将提示您执行"开始您的付费计划"。您可以在个人资料中的订阅 -> 选择您订阅的 API -> 定价标签下找到"开始您的付费计划"按钮。
付款订单在每月 20 日至 30 日之间处理。如果您在 20 日之前提交请求,您的付款将在此时间范围内处理。
您可以通过我们的聊天渠道联系我们以获得即时帮助。我们始终在线,时间为上午 8 点至下午 5 点(EST)。如果您在该时间之后联系我们,我们将尽快回复您。此外,您可以通过 [email protected]

相关 API