I. INTRODUCTION
WebMagic is a simple and flexible Java framework reptiles. Based WebMagic, you can quickly develop a highly efficient, easy to maintain crawlers.
Second, how to learn
1. Check the official website
Official website address is: http://webmagic.io/
official website detailed documentation: http://webmagic.io/docs/zh/
2. Running through hello world example (refer to the official website specifically, reference may be blog)
I wrote the following unit test cases, as a Hello World example.
Maven attention to the need to import dependence:
<dependency> <groupId>us.codecraft</groupId> <artifactId>webmagic-core</artifactId> <version>0.7.3</version> </dependency> <dependency> <groupId>us.codecraft</groupId> <artifactId>webmagic-extension</artifactId> <version>0.7.3</version> </dependency>
3. With an object
Talk about my purpose, I recently developed a blog system, which has a blog to import third-party plug-ins, this is a simple plug-in search box, fill in the corresponding URL inside the search box, click Search into your own blog .
To import blog park single article as an example:
Here is my source code (a single article introduction, I have packaged into one tool class):
import cn.hutool.core.date.DateUtil; import com.blog.springboot.dto.CnBlogModelDTO; import com.blog.springboot.entity.Posts; import com.blog.springboot.service.PostsService; import org.springframework.beans.factory.annotation.Autowired; import org.springframework.stereotype.Component; import us.codecraft.webmagic.Page; import us.codecraft.webmagic.Site; import us.codecraft.webmagic.Spider; import us.codecraft.webmagic.pipeline.ConsolePipeline; import us.codecraft.webmagic.processor.PageProcessor; import us.codecraft.webmagic.selector.Selectable; import javax.annotation.PostConstruct; /** * 导入博客园文章工具类 */ @Component public class WebMagicCnBlogUtils implements PageProcessor { @Autowired private PostsService postService; public static WebMagicCnBlogUtils magicCnBlogUtils; @PostConstruct public void init() { magicCnBlogUtils = this; magicCnBlogUtils.postService = this.postService; } private Site site = Site.me() .setDomain("https://www.cnblogs.com/") .setSleepTime(1000) .setUserAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/67.0.3396.99 Safari/537.36"); @Override public void process(Page page) { Selectable obj = page.getHtml().xpath("//div[@class='post']"); Selectable title = obj.xpath("//h1[@class='postTitle']//a"); Selectable content = obj.xpath("//div[@class='blogpost-body']"); System.out.println("title:" + title.replace("<[^>]*>", "")); System.out.println("content:" + content); CnBlogModelDTO blog = new CnBlogModelDTO(); blog.setTitle(title.toString()); blog.setContent(content.toString()); Posts post = new Posts(); String date = DateUtil.date().toString(); post.setPostAuthor(1L); post.setPostTitle(title.replace("<[^>]*>", "").toString()); post.setPostContent(content.toString()); post.setPostExcerpt(content.replace("<[^>]*>", "").toString()); post.setPostDate(date); post.setPostDate(date); post.setPostModified(date); boolean importPost = magicCnBlogUtils.postService.insert(post); if (importPost) { System.out.println("success"); } else { System.out.println("fail"); } } @Override public Site GetSite () { return Site; } / * * * Importing a single article blog article data Park * * @param url * / public static void importSinglePost (String url) { Spider.create ( new new WebMagicCnBlogUtils ( )) .addUrl (URL) .addPipeline ( new new ConsolePipeline ()) .run (); } }
Unit test code:
import com.blog.springboot.dto.CnBlogModelDTO; import us.codecraft.webmagic.Page; import us.codecraft.webmagic.Site; import us.codecraft.webmagic.Spider; import us.codecraft.webmagic.pipeline.ConsolePipeline; import us.codecraft.webmagic.processor.PageProcessor; import us.codecraft.webmagic.selector.Selectable; public class WebMagicJunitTest implements PageProcessor { private Site site = Site.me() .setDomain("https://www.cnblogs.com/") .setSleepTime(1000) .setUserAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/67.0.3396.99 Safari/537.36"); @Override public void process(Page page) { Selectable obj = page.getHtml().xpath("//div[@class='post']"); Selectable title = obj.xpath("//h1[@class='postTitle']//a"); Selectable content = obj.xpath("//div[@class='blogpost-body']"); System.out.println("title:" + title.replace("<[^>]*>", "")); System.out.println("content:" + content); } @Override public Site getSite() { return site; } public static void importSinglePost(String url) { Spider.create(new WebMagicJunitTest()) .addUrl(url) .addPipeline(new ConsolePipeline()) .run(); } public static void main(String[] args) { WebMagicJunitTest.importSinglePost("https://www.cnblogs.com/youcong/p/9404007.html"); }
In addition, I know what data is how to crawl it?
Needs first, and then by Chrome or Firefox browser checks elements as: