파이썬 - 웹 페이지 데이터 수집을 위한 scrapy Crawler 사용법 요약

윈도우 사용자라고 해도, 간단한 실습에 불과하니 쓸데없이 ^^ 디렉터리를 어지럽히지 말고 WSL에 맡기면 좋습니다. 게다가 파이썬 환경이니 virtualenv로 한 번 더 격리를 하면 좋겠지요. ^^

$ cd ~
~$ mkdir pyenv
~$ cd pyenv

~/pyenv$ virtualenv scraptest

~/pyenv$ source scraptest/bin/activate

(scraptest) testusr@TESTPC:~/pyenv$

scrapy를 설치하고,

(scraptest) testusr@TESTPC:~/pyenv$ pip install scrapy

(scraptest) testusr@TESTPC:~/pyenv$ python --version
Python 3.8.10


(scraptest) testusr@TESTPC:~/pyenv$ scrapy --version
Scrapy 2.5.0 - no active project

Usage:
  scrapy <command> [options] [args]

Available commands:
  bench         Run quick benchmark test
  commands
  fetch         Fetch a URL using the Scrapy downloader
  genspider     Generate new spider using pre-defined templates
  runspider     Run a self-contained spider (without creating a project)
  settings      Get settings values
  shell         Interactive scraping console
  startproject  Create new project
  version       Print Scrapy version
  view          Open URL in browser, as seen by Scrapy

  [ more ]      More commands available when run from project directory

Use "scrapy <command> -h" to see more info about a command

프로젝트를 하나 만듭니다.

(scraptest) testusr@TESTPC:~/pyenv$ cd scraptest/
(scraptest) testusr@TESTPC:~/pyenv/scraptest$

(scraptest) testusr@TESTPC:~/pyenv/scraptest$ scrapy startproject sample1
New Scrapy project 'sample1', using template directory '/home/testusr/pyenv/scraptest/lib/python3.8/site-packages/scrapy/templates/project', created in:
    /home/testusr/pyenv/scraptest/sample1

You can start your first spider with:
    cd sample1
    scrapy genspider example example.com

(scraptest) testusr@TESTPC:~/pyenv/scraptest$ cd sample1/
(scraptest) testusr@TESTPC:~/pyenv/scraptest/sample1$

스파이더를 생성하고,

(scraptest) testusr@TESTPC:~/pyenv/scraptest/sample1$ scrapy genspider sysnet sysnet.pe.kr
Created spider 'sysnet' using template 'basic' in module:
  sample1.spiders.sysnet

(scraptest) testusr@TESTPC:~/pyenv/scraptest/sample1$ tree
.
├── sample1
│   ├── __init__.py
│   ├── __pycache__
│   │   ├── __init__.cpython-38.pyc
│   │   └── settings.cpython-38.pyc
│   ├── items.py
│   ├── middlewares.py
│   ├── pipelines.py
│   ├── settings.py
│   └── spiders
│       ├── __init__.py
│       ├── __pycache__
│       │   └── __init__.cpython-38.pyc
│       └── sysnet.py
└── scrapy.cfg

(scraptest) testusr@TESTPC:~/pyenv/scraptest/sample1$ cat sample1/spiders/sysnet.py
import scrapy

class SysnetSpider(scrapy.Spider):
    name = 'sysnet'
    allowed_domains = ['sysnet.pe.kr']
    start_urls = ['http://sysnet.pe.kr/']

    def parse(self, response):
        pass

여기서, scrap 관련한 코드를 원하는 목적에 맞게 수정해야 합니다. 가령 제 경우에는 제 웹 사이트에 있는 글의 "제목"을 스크립하고 싶은데요, 다음과 같은 식으로 위의 코드를 변경해야 합니다.

# items.py의 기본 내용을 다음과 같이 변경
import scrapy

class Sample1Item(scrapy.Item):
    title = scrapy.Field()

# ./spiders/sysnet.py의 기본 내용을 다음과 같이 변경

import scrapy
from ..items import Sample1Item

class SysnetSpider(scrapy.Spider):
    name = 'sysnet'
    allowed_domains = ['www.sysnet.pe.kr']
    start_urls = ['https://www.sysnet.pe.kr/']

    def start_requests(self):
        maxPage = 5

        for i in range(0, maxPage):
            yield scrapy.Request(self.start_urls[0] + "Default.aspx?mode=2&sub=0&pageno={0}".format(i), self.parse)

    def parse(self, response):
        items = []

        # CSS 예제 (1)
        # for article in response.css('#contentPane > table.postlist > tr'):
        #     item = Sample1Item()
        #     item['title'] = article.css("td:nth-child(5) > a::text").extract()[0]
        #     items.append(item)

        # CSS 예제 (1)
        # for article in response.css('#contentPane > table.postlist > tr > td:nth-child(5) > a::text'):
        #     item = Sample1Item()
        #     item['title'] = article.extract()
        #     items.append(item)

        # XPath 예제 (1)
        # for article in response.xpath('//*[@id="contentPane"]/table[2]/tr'):
        #     item = Sample1Item()
        #     item['title'] = article.xpath("td[5]/a/node()").extract()[0]
        #     items.append(item)

        # XPath 예제 (2)
        for article in response.xpath('//*[@id="contentPane"]/table[2]/tr/td[5]/a/node()'):
            item = Sample1Item()
            item['title'] = article.extract()
            items.append(item)

            print(item['title'])

        return items

# pipelines.py의 기본 내용을 다음과 같이 변경

from itemadapter import ItemAdapter
from scrapy.exporters import JsonItemExporter
import codecs

class Sample1Pipeline:
    def process_item(self, item, spider):
        return item

class JsonPipeline:
    def __init__(self):
        self.file = open('result.json', 'wb')
        self.exporter = JsonItemExporter(self.file, encoding='utf-8', ensure_ascii=False)
        self.exporter.start_exporting()

    def process_item(self, item, spider):
        self.exporter.export_item(item)
        return item

    def close_spider(self, spider):
        self.exporter.finish_exporting()
        self.file.close()

# settings.py에 ITEM_PIPELINES 설정을 새롭게 추가

# ...[생략]...

ITEM_PIPELINES = {
    'sample1.pipelines.JsonPipeline': 300,
}

# ...[생략]...

이렇게 하고, scrapy를 다음과 같이 실행하면,

(scraptest) testusr@TESTPC:~/pyenv/scraptest/sample1$ scrapy crawl sysnet

오류 없이 실행되었으면 result.json 파일에 다음과 같이 5페이지에 해당하는 분량의 제목이 추출된 것을 볼 수 있습니다. ^^

[{"title": "기부/후원"},{"title": ".NET Framework: 1113. C# 10 - (13) 문자열 보간 성능 개선"},{"title": "개발 환경 구성: 603. GoLand - WSL 환경과 연동"},...[생략]...,{"title": ".NET Framework: 1080. xUnit 단위 테스트에 메서드/클래스 수준의 문맥 제공 - Fixture"},{"title": ".NET Framework: 1079. MSTestv2 단위 테스트에 메서드/클래스/어셈블리 수준의 문맥 제공"},{"title": ".NET Framework: 1078. C# 단위 테스트 - MSTestv2/NUnit의 Assert.Inconclusive 사용법(?)"}]

위에서, Spider의 parse 함수 내의 코드에서 결과물을 XPath 또는 CSS를 이용해 select하는 것이 어려울 때는 scrapy를 명령행에서 실행시켜 테스트하는 것이 더 편합니다.

일례로, 위의 경우 페이지 하나를 scrap 하는 다음의 명령을 실행하고,

$ scrapy shell 'https://www.sysnet.pe.kr/Default.aspx?mode=2&sub=0&pageno=0'
...[생략]...
021-09-04 16:56:06 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://www.sysnet.pe.kr/Default.aspx?mode=2&sub=0&pageno=0> (referer: None)
[s] Available Scrapy objects:
[s]   scrapy     scrapy module (contains scrapy.Request, scrapy.Selector, etc)
[s]   crawler    <scrapy.crawler.Crawler object at 0x7f063aad5100>
[s]   item       {}
[s]   request    <GET https://www.sysnet.pe.kr/Default.aspx?mode=2&sub=0&pageno=0>
[s]   response   <200 https://www.sysnet.pe.kr/Default.aspx?mode=2&sub=0&pageno=0>
[s]   settings   <scrapy.settings.Settings object at 0x7f063aad2d90>
[s]   spider     <SysnetSpider 'sysnet' at 0x7f063a7a4100>
[s] Useful shortcuts:
[s]   fetch(url[, redirect=True]) Fetch URL and update local objects (by default, redirects are followed)
[s]   fetch(req)                  Fetch a scrapy.Request and update local objects
[s]   shelp()           Shell help (print this help)
[s]   view(response)    View response in a browser
>>>

진입한 shell 모드에서 다음과 같은 식으로 테스트해 볼 수 있습니다.

>>> response.css('#contentPane > table.postlist > tbody').getall()
[]
>>> response.css('#contentPane > table.postlist > tr').getall()
[...[생략]...]

[이 글에 대해서 여러분들과 의견을 공유하고 싶습니다. 틀리거나 미흡한 부분 또는 의문 사항이 있으시면 언제든 댓글 남겨주십시오.]

[다음 글] .NET Framework: 1114. C# 10 - (13) 단일 파일 내에 적용되는 namespace 선언
[이전 글] .NET Framework: 1113. C# 10 - (12) 문자열 보간 성능 개선

[최초 등록일: 9/4/2021]
[최종 수정일: 10/1/2021]

이 저작물은 크리에이티브 커먼즈 코리아 저작자표시-비영리-변경금지 2.0 대한민국 라이센스에 따라 이용하실 수 있습니다.

by SeongTae Jeong, mailto:techsharer at outlook.com

No	Writer	Date	Cnt.	Title	File(s)
13913	정성태	4/10/2025	6333	오류 유형: 950. Process Explorer - 64비트 윈도우에서 32비트 프로세스의 덤프를 뜰 때 "Error writing dump file: Access is denied." 오류
13912	정성태	4/9/2025	5791	닷넷: 2330. C# - 실행 시에 메서드 가로채기 (.NET 5 ~ .NET 8)	1
13911	정성태	4/8/2025	6904	오류 유형: 949. WinDbg - .NET Core/5+ 응용 프로그램 디버깅 시 sos 확장을 자동으로 로드하지 못하는 문제
13910	정성태	4/8/2025	6121	디버깅 기술: 219. WinDbg - 명령어 내에서 환경 변수 사용법
13909	정성태	4/7/2025	8604	닷넷: 2329. C# - 실행 시에 메서드 가로채기 (.NET Framework 4.8)	1
13908	정성태	4/2/2025	9105	닷넷: 2328. C# - MailKit: SMTP, POP3, IMAP 지원 라이브러리
13907	정성태	3/29/2025	9272	VS.NET IDE: 198. (OneDrive, Dropbox 등의 공유 디렉터리에 있는) C# 프로젝트의 출력 경로 변경하기
13906	정성태	3/27/2025	8734	닷넷: 2327. C# - 초기화되지 않은 메모리에 접근하는 버그?	1
13905	정성태	3/26/2025	8985	Windows: 281. C++ - Windows / Critical Section의 안정화를 위해 도입된 "Keyed Event"	1
13904	정성태	3/25/2025	8212	디버깅 기술: 218. Windbg로 살펴보는 Win32 Critical Section	1
13903	정성태	3/24/2025	8201	VS.NET IDE: 197. (OneDrive, Dropbox 등의 공유 디렉터리에 있는) C++ 프로젝트의 출력 경로 변경하기
13902	정성태	3/24/2025	7937	개발 환경 구성: 742. Oracle - 테스트용 hr 계정 및 데이터 생성	1
13901	정성태	3/9/2025	8295	Windows: 280. Hyper-V의 3가지 Thread Scheduler (Classic, Core, Root)
13900	정성태	3/8/2025	10476	스크립트: 72. 파이썬 - SQLAlchemy + oracledb 연동
13899	정성태	3/7/2025	7046	스크립트: 71. 파이썬 - asyncio의 ContextVar 전달
13898	정성태	3/5/2025	7551	오류 유형: 948. Visual Studio - Proxy Authentication Required: dotnetfeed.blob.core.windows.net
13897	정성태	3/5/2025	9466	닷넷: 2326. C# - PowerShell과 연동하는 방법 (두 번째 이야기)	1
13896	정성태	3/5/2025	9501	Windows: 279. Hyper-V Manager - VM 목록의 CPU Usage 항목이 항상 0%로 나오는 문제
13895	정성태	3/4/2025	9132	Linux: 117. eBPF (bpf2go) - Map에 추가된 요소의 개수를 확인하는 방법
13894	정성태	2/28/2025	8147	Linux: 116. eBPF (bpf2go) - BTF Style Maps 정의 구문과 데이터 정렬 문제
13893	정성태	2/27/2025	7396	Linux: 115. eBPF (bpf2go) - ARRAY / HASH map 기본 사용법
13892	정성태	2/24/2025	10683	닷넷: 2325. C# - PowerShell과 연동하는 방법	1
13891	정성태	2/23/2025	8077	닷넷: 2324. C# - 프로세스의 성능 카운터용 인스턴스 이름을 구하는 방법	1
13890	정성태	2/21/2025	9236	닷넷: 2323. C# - 프로세스 메모리 중 Private Working Set 크기를 구하는 방법(Win32 API)	1
13889	정성태	2/20/2025	10100	닷넷: 2322. C# - 프로세스 메모리 중 Private Working Set 크기를 구하는 방법(성능 카운터, WMI) [1]	1
13888	정성태	2/17/2025	10655	닷넷: 2321. Blazor에서 발생할 수 있는 async void 메서드의 부작용

AD BLOCK 해제 요청

파이썬 - 웹 페이지 데이터 수집을 위한 scrapy Crawler 사용법 요약