首页 > temp > python入门教程 >
-
Python爬虫一键下载yy全站短视频详细步骤(附源码)
基本开发环境
- Python 3.6
- Pycharm
相关模块的使用
import os import requests
安装Python并添加到环境变量,pip安装需要的相关模块即可。
一、确定目标需求
data:image/s3,"s3://crabby-images/642dc/642dc0816ee702de010254332cef188f2a0593e9" alt=""
data:image/s3,"s3://crabby-images/a3fa1/a3fa1cbd386856b3073022090aa923c3357a4aaf" alt=""
百度搜索YY,点击分类选择小视频,里面的小姐姐自拍的短视频就是我们所需要的数据了。
data:image/s3,"s3://crabby-images/3d0bc/3d0bc14da507c2611cea5bec3a0f8cc9a9d6d6df" alt=""
二、网页数据分析
网站是下滑网页之后加载数据,在上篇关于好看视频的爬取文章中已经有说明,YY视频也是换汤不换药。
data:image/s3,"s3://crabby-images/a3829/a3829d4b3aa59f5a1febcb2be081891496787b33" alt=""
如图所示,所框选的url地址,就是短视频的播放地址了。
data:image/s3,"s3://crabby-images/f0113/f01130c19d5d9175e2a42e4c96712c66c1bca8dd" alt=""
数据包接口地址:
https://api-tinyvideo-web.yy.com/home/tinyvideosv2?callback=jQuery112409962628943012035_1613628479734&appId=svwebpc&sign=&data=%7B%22uid%22%3A0%2C%22page%22%3A1%2C%22pageSize%22%3A10%7D&_=1613628479736
第二页的数据请求参数:
data:image/s3,"s3://crabby-images/3fb7b/3fb7b62df221ad9bc7863b7caab3469701382425" alt=""
第三页的数据请求参数:
data:image/s3,"s3://crabby-images/84e89/84e89b3c24a7c77b093a016abda0bd1cb6e4b2b7" alt=""
很明显这是根据data参数中的page改变翻页的。
构建翻页循环,获取视频url地址以及发布人的名字,保存到本地。
三、代码实现
1、请求数据接口
import requests url = 'https://api-tinyvideo-web.yy.com/home/tinyvideosv2' params = { 'callback': 'jQuery112409962628943012035_1613628479734', 'appId': 'svwebpc', 'sign': '', 'data': '{"uid":0,"page":0,"pageSize":10}', '_': '1613628479737', } headers = { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.138 Safari/537.36' } response = requests.get(url=url, params=params, headers=headers)
问题来了,返回的数据是json数据嘛?
data:image/s3,"s3://crabby-images/18cdb/18cdb3439ebe3294f35fdf9ba8ab422413c90dc1" alt=""
如上图所示,很多人看到这样的数据肯定就觉得这不就是一个json数据嘛?
data:image/s3,"s3://crabby-images/8b575/8b575f26987d6091f6f378e5a37c47d08edd9492" alt=""
JSONDecodeError: json解码错误,它并不是有一个json数据,而是字符串。
data:image/s3,"s3://crabby-images/50ea4/50ea4d754d5298ce5635e679eea3f06a391dd3cb" alt=""
通过response查看就知道了,返回给我们的数据是多了一段 jQuery112409962628943012035_1613628479734()
其中的json数据是包含在里面的,如果想要提取数据有三种方法。
1、返回response.text,使用正则表达式提取url地址以及发布人的名字
video_url = re.findall('"resurl":"(.*?)"', response.text) user_name = re.findall('"username":"(.*?)"', response.text)
data:image/s3,"s3://crabby-images/2fade/2fadec52ea87dadd7f4e66a00729d10a15657d6e" alt=""
2、返回response.text,使用正则表达式提取 jQuery112409962628943012035_1613628479734() 中的数据,然后通过json模块把字符串转成json数据,然后遍历提取数据。
string = re.findall('jQuery112409962628943012035_1613628479734\((.*?)\)', response.text)[0] json_data = json.loads(string) result = json_data['data']['data'] pprint.pprint(result)
data:image/s3,"s3://crabby-images/9e4b7/9e4b78e6050af2d57db4d5fc83cd22c296b639f8" alt=""
3、把请求的url地址中的 callback 删掉,可以直接获取json数据
import pprint import requests url = 'https://api-tinyvideo-web.yy.com/home/tinyvideosv2' params = { 'appId': 'svwebpc', 'sign': '', 'data': '{"uid":0,"page":1,"pageSize":10}', '_': '1613628479737', } headers = { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.138 Safari/537.36' } response = requests.get(url=url, params=params, headers=headers) json_data = response.json() result = json_data['data']['data'] pprint.pprint(result)
2、保存数据
for index in result: video_url = index['resurl'] user_name = index['username'] video_content = requests.get(url=video_url, headers=headers).content with open('video\\' + user_name + '.mp4', mode='wb') as f: f.write(video_content) print(user_name)
注意点: 用户名有特殊字符,保存的时候会报错
data:image/s3,"s3://crabby-images/4a117/4a1178bd71bb77a51c26da514218033caf963de5" alt=""
data:image/s3,"s3://crabby-images/0a0d5/0a0d5d2617319a818a89f1d7f44ad54244925143" alt=""
所以需要使用正则表达式替换掉特殊字符
def change_title(title): pattern = re.compile(r"[\/\\\:\*\?\"\<\>\|]") # '/ \ : * ? " < > |' new_title = re.sub(pattern, "_", title) # 替换为下划线 return new_title
完整实现代码
import re import requests import re def change_title(title): pattern = re.compile(r"[\/\\\:\*\?\"\<\>\|]") # '/ \ : * ? " < > |' new_title = re.sub(pattern, "_", title) # 替换为下划线 return new_title page = 0 while True: page += 1 url = 'https://api-tinyvideo-web.yy.com/home/tinyvideosv2' params = { 'appId': 'svwebpc', 'sign': '', 'data': '{"uid":0,"page":%s,"pageSize":10}' % str(page), '_': '1613628479737', } headers = { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.138 Safari/537.36' } response = requests.get(url=url, params=params, headers=headers) json_data = response.json() result = json_data['data']['data'] for index in result: video_url = index['resurl'] user_name = index['username'] new_title = change_title(user_name) video_content = requests.get(url=video_url, headers=headers).content with open('video\\' + new_title + '.mp4', mode='wb') as f: f.write(video_content) print(user_name)
文章出处:https://www.cnblogs.com/hhh188764/p/14417773.html
data:image/s3,"s3://crabby-images/8fe46/8fe4613f931bed234d537c6cf031c8806d0da0e9" alt=""