<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Sergio Segura</title>
    <description>Informática, desarrollo y reflexiones</description>
    <link>https://suresrm.github.io//</link>
    <atom:link href="https://suresrm.github.io//feed.xml" rel="self" type="application/rss+xml" />
    
      <item>
        <title>Web Scraping for Data Analysis in Python (Parte V): Enriquecimiento de datos</title>
        <description>&lt;p&gt;En &lt;a href=&quot;/2019/11/22/geoencoding/&quot;&gt;el artículo anterior&lt;/a&gt; conseguimos información espacial de cada uno de los items, algo que ya adelantamos que nos abriría muchas posibilidades en el futuro. En este artículo haremos uso de dicha información para cruzarla con otras fuentes de datos y así obtener información más relevante.&lt;/p&gt;

&lt;p&gt;Existen diversas formas de realizar las operaciones espaciales que vamos a efectuar. El criterio que se ha seguido para elegir una herramienta u otra depende exclusivamente de la experiencia previa que se tenga con la misma. En este caso particular, he elegido utilizar la base de datos relacional con soporte para operaciones espaciales &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;spatialite&lt;/code&gt;.&lt;/p&gt;

&lt;h2 id=&quot;configuración&quot;&gt;Configuración&lt;/h2&gt;

&lt;p&gt;Una de las grandes ventajas de &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;spatialite&lt;/code&gt; es su portabilidad. Se trata de una herramienta auto-contenida que puede ser ejecutada desde cualquier directorio. El paquete se encuentra en los repositorios de Debian y Ubuntu por que puede ser instalada mediante &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sudo apt install spatialite-gui&lt;/code&gt;. Esto instalará la librería junto con su entorno gráfico.&lt;/p&gt;

&lt;p&gt;Una vez instalado se puede iniciar ejecutando &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;spatialite-gui&lt;/code&gt; en la terminal. Esto nos abrirá una ventana con la interfaz gráfica.&lt;/p&gt;

&lt;h2 id=&quot;carga-de-datos&quot;&gt;Carga de datos&lt;/h2&gt;

&lt;p&gt;Los datos que queremos cruzar en este caso de estudio provienen del Instituto Nacional de Estadística y expresan el salario medio por hogar con un grano de de detalle a nivel de sección administrativa (subdivisiones inferiores a un barrio). El problema es que dichos datos no contienen información espacial alguna, se trata de una simple tabla de tres columnas: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sección&lt;/code&gt; , &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;salario-por-hogar&lt;/code&gt; y &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;salario-por-persona&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Para atajar esto usaremos una tercera fuente de datos, una capa espacial que contiene las geometrías de dichas secciones y nos permitirá determinar a qué sección pertenece cada item.&lt;/p&gt;

&lt;p&gt;En nuestra base de datos tendremos, entonces:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Una capa de puntos con los items (pisos)&lt;/li&gt;
  &lt;li&gt;Una capa de geometrías con los distritos&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 id=&quot;carga-de-los-distritos&quot;&gt;Carga de los distritos&lt;/h3&gt;

&lt;p&gt;Para cargar la capa de geometrías, importaremos el dataset descargado. Al tratarse de un dataset en formato Shapefile, podemos importarlo directamente clickando en el icono &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Load Shapefile&lt;/code&gt;. Si estuviera en otro formato no compatible, sería necesario buscar la forma de convertirlo a Shapefile u a otro formato compatible.&lt;/p&gt;

&lt;p&gt;En el proceso de importación se requerirá elegir el nombre de la tabla destino (en nuestro caso &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;distritos&lt;/code&gt;) y el nombre de la columna espacial (la llamaremos &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;geom&lt;/code&gt;) así como otros parámetros específicos como el SRID. El SRID indica el sistema de proyección usado y debería estar documentado en la fuente del dataset o en algún metadato adjunto. En el caso de los datos de distritos obtenidos, su SRID es &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;3042&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Si no fuera posible averiguar el SRID, se puede probar a utilizar software especializado como Quantum GIS, mucho más potente y versatil, pero con una complejidad que escapa a este artículo. Otra ventaja de incluir este software en el proceso será la capacidad de visualizar los resultados proyectados sobre un mapa.&lt;/p&gt;

&lt;h3 id=&quot;carga-de-los-items&quot;&gt;Carga de los items&lt;/h3&gt;

&lt;p&gt;La carga de los item se hará con otra de las opciones compatibles que ofrece &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;spatialite-gui&lt;/code&gt;, la carga desde CSV. Esto se logra clickando en &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Load CSV/TXT&lt;/code&gt;, eligiendo el fichero CSV correspondiente y completando los datos de formato del mismo, como el separador usado (habitualmente coma o tabulador), el nombre (en nuestro caso &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;points&lt;/code&gt;) o el SRID. En este caso, no encontraremos documentación alguna sobre el SRID ya que estos datos nos hemos generado nosotros. No obstante, sabemos que los datos &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lat&lt;/code&gt; y &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lon&lt;/code&gt; provienen de coordenadas globales, a las cuales les corresponde el SRID &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;4326&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;La tabla creada carecerá de una columna espacial por lo que deberemos crearla ejecutando las siguientes dos intrucciones desde la consola SQL:&lt;/p&gt;

&lt;div class=&quot;language-sql highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;SELECT&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;AddGeometryColumn&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;points&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;geom&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;4326&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;POINT&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;UPDATE&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;points&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;SET&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;geom&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;MakePoint&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;lon&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;lat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;4326&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Esto debería resultar en la creación de una columna nueva llamada &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;geom&lt;/code&gt; con los correspondientes puntos en formato espacial. Podemos comprobar el proceso expandiendo la tabla &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;points&lt;/code&gt; en la vista en árbol hasta ver el atributo &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;geom&lt;/code&gt;, hacer click sobre él y seleccionar &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Map Preview&lt;/code&gt;. Esta previsualización debería mostrar los puntos correspondientes a cada item.&lt;/p&gt;

&lt;h2 id=&quot;consulta-espacial&quot;&gt;Consulta espacial&lt;/h2&gt;

&lt;p&gt;El siguiente paso es realizar lo que en tratamiento de datos de conoce como &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lookup&lt;/code&gt;. Básicamente, trataremos de añadir a cada item el distrito en el que se encuentra buscando la geometría en la que se encuentra contenido. Para ello utilizaremos la función de &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;spatialite&lt;/code&gt; &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;within()&lt;/code&gt; que nos dice si una geometría A se encuentra contenida en una geometría B.&lt;/p&gt;

&lt;div class=&quot;language-sql highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;select&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;title&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cusec&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;points&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;distritos&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;d&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;where&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;within&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ST_Transform&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;p&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;geom&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;3042&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;geom&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;AND&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;d&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;cpro&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;50&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;De este modo, el resultado es una lista de tres columnas con el &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;id&lt;/code&gt;, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;titulo&lt;/code&gt; y &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;distrito&lt;/code&gt; de cada item. El código de &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;distrito&lt;/code&gt; será el que podremos usar para cruzar con cualquier otro dato externo, como la renta media que antes mencionábamos.&lt;/p&gt;

&lt;p&gt;Para cruzar las tres tablas: &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;points&lt;/code&gt;, el resultado con los códigos de distrito y la fuente de datos externa, se puede usar la propia base de datos mediante SQL o se puede exportar el resultado a CSV haciendo click derecho sobre él y posteriormente agregar todo en una hoja de cálculo como Libre Office Calc o Excel.&lt;/p&gt;

&lt;h2 id=&quot;distancias&quot;&gt;Distancias&lt;/h2&gt;

&lt;p&gt;Una de las ventajas de tener los datos cargados en la base de datos espacial es que, con poco esfuerzo extra, podemos lograr extraer datos muy interesantes. Por ejemplo distancias mínimas (reales geográficas, no aproximaciones euclídeas) a puntos, trazados o polígonos. En nuestro caso de estudio, ya que se está recolectando una base de datos de pisos de alquiler de interés, puede resultar interesante conocer la distancia tu centro de trabajo, a la boca de metro más cercana o al tranvía que atraviesa tu cuidad.&lt;/p&gt;

&lt;p&gt;Para ello, la consulta es muy similar a la anterior en estructura, pero variando la función utilizada. Además, ya no cruzaremos dos geometrías existentes, sino que proveeremos una de ellas.&lt;/p&gt;

&lt;div class=&quot;language-sql highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;select&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;id&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;ST_Distance&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
		&lt;span class=&quot;n&quot;&gt;ST_GeomFromText&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
			&lt;span class=&quot;s1&quot;&gt;&apos;LINESTRING(-0.8708779781997009 41.68653240939767,-0.8898465603530212 41.687109287806656,-0.8900182217299744 41.679096624709025,-0.8901898831069275 41.6703135999149,-0.8838586602971645 41.66353877155025,-0.8843736444280239 41.661230381694274,-0.881026247577438 41.66090976544719,-0.8815841470525356 41.65776764174951,-0.8842878137395473 41.65404819516039,-0.8805541787908169 41.652124259193606,-0.8857898507878872 41.64728209938806,-0.8910513469900252 41.644907918854656,-0.898046548100865 41.63743525238525,-0.905284253258742 41.63226155866651,-0.9126274231276739 41.62269179353281,-0.9201805237136114 41.615986573308966,-0.9274257600052351 41.619661240493386)&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
			&lt;span class=&quot;mi&quot;&gt;4326&lt;/span&gt;
		&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
		&lt;span class=&quot;n&quot;&gt;geom&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;
	&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;dist_tranv&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;ST_Distance&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;
		&lt;span class=&quot;n&quot;&gt;ST_GeomFromText&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s1&quot;&gt;&apos;POINT(-0.8892394427357431 41.68358071844347)&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;4326&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;
		&lt;span class=&quot;n&quot;&gt;geom&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;1&lt;/span&gt;
	&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;dist_cps&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;points&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Las geometrías o los puntos pueden generarse desde cualquier editor on-line de WKT (este es el formato de texto en el que definimos los puntos o las líneas).&lt;/p&gt;

&lt;p&gt;Igual que en el caso anterior, el resultado puede ser agregado en la propia base de datos o en una hola de cálculo, la decisión depende de las preferencias y conocimientos de cada uno.&lt;/p&gt;
</description>
        <pubDate>Sat, 30 Nov 2019 00:00:00 +0000</pubDate>
        <link>https://suresrm.github.io//2019/11/30/enriquecimiento/</link>
        <guid isPermaLink="true">https://suresrm.github.io//2019/11/30/enriquecimiento/</guid>
      </item>
    
      <item>
        <title>Web Scraping for Data Analysis in Python (Parte IV): Geoencoding</title>
        <description>&lt;p&gt;En &lt;a href=&quot;/2019/11/15/scrapping/&quot;&gt;el artículo anterior&lt;/a&gt; obtuvimos todos los datos estructurados que el portal web nos podía ofrecer. No obstante, podemos ir mucho (mucho) más allá aprovechando el resto de la web como fuente de datos. Una de las características más importantes que podemos desear conocer es la posición espacial que ocupa el item. Conocer las coordenadas exactas (o aproximadas) de un dato nos abre la puerta a un océano de datos espaciales con los que cruzar nuestros datos.&lt;/p&gt;

&lt;p&gt;En nuestro caso de estudio, los pisos no cuentan con dicha información de forma fácilmente extraíble, pero si conocemos el nombre de la calle o la zona. Esto nos permite aplicar una técnica conocida como geoencoding: traducir de nombres a lugares.&lt;/p&gt;

&lt;h2 id=&quot;presentando-nomatim-by-osm&quot;&gt;Presentando: Nomatim (by OSM)&lt;/h2&gt;

&lt;p&gt;&lt;a href=&quot;https://nominatim.openstreetmap.org/search.php?q=plaza+para%C3%ADso&quot;&gt;Nomatim&lt;/a&gt; es un servicio abierto de geoencoding basado en Open Street Maps. El servicio ofrece búsquedas ilimitadas y gratuitas siempre que se respeten unos acuerdos de uso, que restringen, entre otras cosas, la frecuencia de las peticiones.&lt;/p&gt;

&lt;h2 id=&quot;configuración&quot;&gt;Configuración&lt;/h2&gt;

&lt;p&gt;Los pre-requisitos en este caso serán:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python==3.6&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;requests==2.22.0&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;La librería &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;requests&lt;/code&gt;, de hecho, no sería estrictamente necesaria, pero facilitará enormemente la realización de las peticiones al servidor.&lt;/p&gt;

&lt;h2 id=&quot;funcionamiento&quot;&gt;Funcionamiento&lt;/h2&gt;

&lt;h3 id=&quot;la-api-de-nomatim&quot;&gt;La API de Nomatim&lt;/h3&gt;

&lt;p&gt;El servicio expone una API sencilla con la que podremos consultar por nombres de calles y plazas siguiendo un formato específico:&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;GET https://nominatim.openstreetmap.org/search?q&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;
	&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;Nombre&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;+De+La+Calle&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;
	&amp;amp;polygon_geojson&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;1
	&amp;amp;viewbox&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;latlon_left&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;%&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;latlon_top&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;%2C&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;latlon_right&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;%2&lt;span class=&quot;k&quot;&gt;${&lt;/span&gt;&lt;span class=&quot;nv&quot;&gt;latlon_bot&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;}&lt;/span&gt;
	&amp;amp;format&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;json
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;El resultado será un objeto JSON (ya que así ha sido solicitado en la API) con varios campos de interés como el &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;display_name&lt;/code&gt; que permite supervisar que el resultado es el esperado, o los campos &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lat&lt;/code&gt; y &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;lon&lt;/code&gt; que, como sus nombres indican, contienen las coordenadas del centroide de la geometría. Esto puede suponer un problema en el caso de calles largas, pero es la mejor aproximación que se puede realizar sin tener los números de los portales.&lt;/p&gt;

&lt;p&gt;El código python para llamar a la API para nuestro caso de estudio:&lt;/p&gt;

&lt;div class=&quot;language-py highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;url&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;https://nominatim.openstreetmap.org/search?q=&apos;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; \
	&lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;strip&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;().&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;replace&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos; &apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;+&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;+Zaragoza+Aragon+Spain&apos;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; \
	&lt;span class=&quot;s&quot;&gt;&apos;&amp;amp;polygon_geojson=1&apos;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; \
	&lt;span class=&quot;s&quot;&gt;&apos;&amp;amp;viewbox=-1.30875%2C41.77592%2C-0.43327%2C41.51937&apos;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;+&lt;/span&gt; \
	&lt;span class=&quot;s&quot;&gt;&apos;&amp;amp;format=json&apos;&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;headers&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
	&lt;span class=&quot;s&quot;&gt;&apos;User-Agent&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;Python Script for Economics Study&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;results&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;requests&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;headers&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;headers&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;json&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;algoritmo&quot;&gt;Algoritmo&lt;/h3&gt;

&lt;p&gt;El algoritmo a ejecutar no entraña misterio alguno:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Cargar el fichero CSV con todos los items&lt;/li&gt;
  &lt;li&gt;Iterar, para cada item 1.
    &lt;ol&gt;
      &lt;li&gt;Extraer su dirección
        &lt;ol&gt;
          &lt;li&gt;“Limpiar” dicha dirección&lt;/li&gt;
          &lt;li&gt;Realizar una búsqueda en nominatim&lt;/li&gt;
          &lt;li&gt;Comprobar si, entre la lista de resultados, hay un resultado válido (dentro de un área razonable, por ejemplo)&lt;/li&gt;
          &lt;li&gt;Si lo hay, continuar con el siguiente item&lt;/li&gt;
          &lt;li&gt;Si no la hay, probar con otra “limpieza” y volver al principio de este bucle&lt;/li&gt;
        &lt;/ol&gt;
      &lt;/li&gt;
    &lt;/ol&gt;
  &lt;/li&gt;
&lt;/ol&gt;

&lt;h3 id=&quot;limpieza-de-datos&quot;&gt;Limpieza de datos&lt;/h3&gt;

&lt;p&gt;El filtrado de nombres es uno de los pasos más artesanales y más guiados por la prueba-error de todo el proceso. Se han detectado que los términos que más empeoran la búsqueda son ‘calle’, ‘plaza’, ‘avenida’ y ‘paseo’. La estrategia es detectar estos términos y eliminarlos del nombre.&lt;/p&gt;

&lt;p&gt;También se han observado problemas con el artículo ‘de’, por lo que otro de los “niveles de limpieza” implica eliminar dichos artículos.&lt;/p&gt;

&lt;p&gt;El código de limpieza de nombres se puede condensar en las siguientes dos funciones.&lt;/p&gt;

&lt;div class=&quot;language-py highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;rip&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
	&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;calle&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;plaza&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;avenida&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;paseo&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]:&lt;/span&gt;
		&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
			&lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;_rip_pattern&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
	&lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;_rip_pattern&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
		&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos; de&apos;&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;not&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
			&lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;replace&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos; de&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
		&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos; de&apos;&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
			&lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;replace&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;pattern&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;+&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos; de&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
		&lt;span class=&quot;k&quot;&gt;return&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;query&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Al final del proceso, se han logrado ubicar cerca del 80% de los pisos de nuestro caso práctico combinando las distintas técnicas de “limpieza”.&lt;/p&gt;
</description>
        <pubDate>Fri, 22 Nov 2019 00:00:00 +0000</pubDate>
        <link>https://suresrm.github.io//2019/11/22/geoencoding/</link>
        <guid isPermaLink="true">https://suresrm.github.io//2019/11/22/geoencoding/</guid>
      </item>
    
      <item>
        <title>Web Scraping for Data Analysis in Python (Parte III): Scrapping</title>
        <description>&lt;p&gt;En &lt;a href=&quot;/2019/11/08/proteccion/&quot;&gt;el artículo anterior&lt;/a&gt; descargamos una copia local de todas las páginas que contenían información relevante para nosotros. Esto nos va a permitir trabajar con libertad a partir de este momento ya que no necesitaremos volver a acceder al portal. En otras palabras, podemos trabajar completamente offline.&lt;/p&gt;

&lt;h2 id=&quot;presentando-scrapy&quot;&gt;Presentando: Scrapy&lt;/h2&gt;

&lt;p&gt;Mientras Selenium era una librería capaz de darnos el control absoluto de un navegador completo, Scrapy sigue una aproximación diferente: Analiza páginas HTML y ofrece un método muy sencillo para extraer información estructurada de ellas en formatos como CSV o JSON.&lt;/p&gt;

&lt;p&gt;En muchos casos, Scrapy puede ser der usada directamente contra el servidor, pero en el caso de nuestro portal, se observó que el funcionamiento natural de la herramienta hacía que fuera muy fácil de detectar y las configuraciones necesarias para alterar ese funcionamiento eran complejas e inconsistentes.&lt;/p&gt;

&lt;h2 id=&quot;configuración&quot;&gt;Configuración&lt;/h2&gt;

&lt;p&gt;De modo similar a Selenium, los pre-requisitos son:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python==3.6&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scrapy==1.7.3&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;En este caso, al no usar un navegador ni ningún otro servicio no será necesario descargar nada más. Por otro lado, sí que habrá que conocer la estructura de programa que Scrapy requiere, ya que éste es mucho más restrictivo que Selenium en ese aspecto.&lt;/p&gt;

&lt;h2 id=&quot;funcionamiento&quot;&gt;Funcionamiento&lt;/h2&gt;

&lt;h3 id=&quot;la-clase-spyder&quot;&gt;La clase Spyder&lt;/h3&gt;

&lt;p&gt;La naturaleza de &lt;em&gt;framework&lt;/em&gt; de Scrapy implica que hemos de adaptarnos a la estructura que propone, en este caso, extendiendo la clase &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Spider&lt;/code&gt; en un fichero python llamado, por ejemplo, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scrapy.py&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Fragmento extraído de la documentación oficial&lt;/em&gt;&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;scrapy&lt;/span&gt;


&lt;span class=&quot;k&quot;&gt;class&lt;/span&gt; &lt;span class=&quot;nc&quot;&gt;MySpyfer&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;scrapy&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Spider&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
    &lt;span class=&quot;n&quot;&gt;name&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&quot;quotes&quot;&lt;/span&gt;

    &lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;start_requests&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;urls&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;
            &lt;span class=&quot;s&quot;&gt;&apos;http://quotes.toscrape.com/page/1/&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
            &lt;span class=&quot;s&quot;&gt;&apos;http://quotes.toscrape.com/page/2/&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
        &lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;url&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;urls&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;k&quot;&gt;yield&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;scrapy&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Request&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;callback&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;parse&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

    &lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;parse&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;response&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;page&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;response&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;url&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;split&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;/&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)[&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;-&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;2&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;
        &lt;span class=&quot;n&quot;&gt;filename&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;quotes-%s.html&apos;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;%&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;page&lt;/span&gt;
        &lt;span class=&quot;k&quot;&gt;with&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;open&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;filename&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;wb&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
            &lt;span class=&quot;n&quot;&gt;f&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;write&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;response&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;body&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
        &lt;span class=&quot;bp&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;log&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;Saved file %s&apos;&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;%&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;filename&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;ul&gt;
  &lt;li&gt;La función &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;start_requests()&lt;/code&gt; debe devolver iterativamente objetos de tipo &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Request&lt;/code&gt; con la dirección deseada.&lt;/li&gt;
  &lt;li&gt;La función &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;parse()&lt;/code&gt; será llamada por Scrapy recibiendo el objeto &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Response&lt;/code&gt; correspondiente a cada &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Request&lt;/code&gt;. De este objeto obtendremos toda la información deseada.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 id=&quot;creación-de-las-uri&quot;&gt;Creación de las URI&lt;/h3&gt;

&lt;p&gt;En nuestro caso, como las páginas estan descargadas localmente, las URI han de tener la estructura:
&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;file:///ruta/local_directorio/fichero.html&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Otra opción es ejecutar un servidor local como &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;http-server&lt;/code&gt; (de Node).
scrapy crawl main -o ./data/main.csv&lt;/p&gt;

&lt;h3 id=&quot;scrapping-de-la-página&quot;&gt;Scrapping de la página&lt;/h3&gt;

&lt;p&gt;Este apartado es fuertemente dependiente del sitio web a &lt;em&gt;scrapear&lt;/em&gt;, ya que todas las directivas &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;css&lt;/code&gt; y &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;xpath&lt;/code&gt; son específicas para portal.&lt;/p&gt;

&lt;p&gt;Tal y como se introdujo en con Selenium, el método de selección de elementos que usaremos será la sintaxis de CSS combinada con la sintaxis XPath (similar, pero más potente en algunos aspectos). El funcionamiento específico de ambas sintaxis es extenso y hay suficiente literatura escrita como para profundizar en ello todo lo que se desee.&lt;/p&gt;

&lt;p&gt;En el caso que nos ocupa, los campos pueden ser extraidos de la siguiente manera:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;parse&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;bp&quot;&gt;self&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;response&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;):&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;flat&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;response&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;css&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;section.detail-info&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

	&lt;span class=&quot;n&quot;&gt;title&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;flat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;css&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;span.main-info__title-main::text&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;

	&lt;span class=&quot;n&quot;&gt;description&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;flat&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;css&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;.adCommentsLanguage&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;xpath&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;.//text()&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;).&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;getall&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;description&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos; &apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;join&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;description&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Los dos casos básicos son, extracción de campos simples, como el título; y extración de listas, como la descripción, que luego es concatenada con espacios.&lt;/p&gt;

&lt;h3 id=&quot;devolución-del-resultado&quot;&gt;Devolución del resultado&lt;/h3&gt;

&lt;p&gt;Finalmente, con todos los campos extraídos en variables, queda sólo devolver el resultado en forma de diccionario python:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;yield&lt;/span&gt; &lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;
	&lt;span class=&quot;s&quot;&gt;&apos;title&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;title&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;
	&lt;span class=&quot;s&quot;&gt;&apos;description&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;description&lt;/span&gt;
&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;ejecución-del-experimento&quot;&gt;Ejecución del experimento&lt;/h2&gt;

&lt;p&gt;Para lanzar el &lt;em&gt;spider&lt;/em&gt;, será necesario ejecutar:&lt;/p&gt;

&lt;div class=&quot;language-sh highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;python &lt;span class=&quot;nt&quot;&gt;-m&lt;/span&gt; scrapy runspider scrapy.py &lt;span class=&quot;nt&quot;&gt;-o&lt;/span&gt; .data.csv
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Esto iniciará el proceso de recolección de datos y debería generar un fichero CSV con una fila por cada item y una columna por cada campo devuelto en el diccionario python.&lt;/p&gt;
</description>
        <pubDate>Fri, 15 Nov 2019 00:00:00 +0000</pubDate>
        <link>https://suresrm.github.io//2019/11/15/scrapping/</link>
        <guid isPermaLink="true">https://suresrm.github.io//2019/11/15/scrapping/</guid>
      </item>
    
      <item>
        <title>Web Scraping for Data Analysis in Python (Parte II): Protección contra crawlers</title>
        <description>&lt;p&gt;Tal y como se mencionó en &lt;a href=&quot;/2019/11/01/introduccion/&quot;&gt;la introducción&lt;/a&gt;, algunos sitios web establecen medidas de seguridad para evitar el acceso de crawlers. Tales medidas pueden abarcar desde una simple comprobación del &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;User-Agent&lt;/code&gt; hasta un sofisticado sistema de registros y heurísticas que tratan de detectar a estos autómatas.&lt;/p&gt;

&lt;p&gt;Nuestro trabajo será, precisamente, engañar a esos sistemas para evitar ser detectados y preparar una estrategia de recuperación en caso de bloqueo.&lt;/p&gt;

&lt;h2 id=&quot;presentando-selenium&quot;&gt;Presentando: Selenium&lt;/h2&gt;

&lt;p&gt;Cuando de trata de automatizar procesos sobre un navegador, Selenium sin duda la herramienta de referencia. Una de sus mayores ventajas es que, a diferencia de otras herramientas como Scrapy, que trabajan sobre documentos HTML descargados estáticamente, Selenium ejecuta y controla un navegador web al completo, por lo que es mucho más difícil de detectar.&lt;/p&gt;

&lt;p&gt;Selenium esta disponible para multitud de lenguajes como Java, Javascript, Python, etc. Para este proyecto se ha decidido utilizar Python a lo largo de todo el proceso, tanto por la disponibilidad de las herramientas como por la popularidad del lenguaje entre sectores &lt;em&gt;no-tic&lt;/em&gt;.&lt;/p&gt;

&lt;h2 id=&quot;configuración&quot;&gt;Configuración&lt;/h2&gt;

&lt;p&gt;Los pre-requisitos para seguir en proceso son:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python==3.6&lt;/code&gt;&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;selenium==3.141.0&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;En el directorio actual, crearemos el script con el nombre que deseemos, por ejemplo, &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;crawl.py&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;Comprobaremos que las dependencias son correctas escribiendo en el script:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;selenium&lt;/span&gt; &lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;webdriver&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;Done&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;De este modo, ejecutando &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;python crawl.py&lt;/code&gt; deberíamos ver escrito por pantalla el mensaje &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;Done&lt;/code&gt;.&lt;/p&gt;

&lt;h3 id=&quot;obtenicón-del-webdriver&quot;&gt;Obtenicón del WebDriver&lt;/h3&gt;

&lt;p&gt;Selenium no trae instalado por defecto ningún navegador, por ello se ha de escoger uno y obtener su WebDriver por separado. El WebDriver es lo que le permitirá tomar el control del navegador, y ha de ser específico para cada navegador distinto. En este caso, optaremos por Firefox al ser el navegador de código libre más popular y robusto. El WebDirver se puede encontrar en &lt;a href=&quot;https://github.com/mozilla/geckodriver/releases&quot;&gt;su repositorio oficial&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Una vez descargado y descomprimido en el directorio local, podemos probar su funcionamiento editando el script.&lt;/p&gt;
&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;kn&quot;&gt;from&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;selenium&lt;/span&gt; &lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;webdriver&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;os&lt;/span&gt;
&lt;span class=&quot;kn&quot;&gt;import&lt;/span&gt; &lt;span class=&quot;nn&quot;&gt;time&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;BASE_PATH&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;abspath&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dirname&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;__file__&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;DRIVER_PATH&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;join&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BASE_PATH&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;geckodriver&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;CACHE_PATH&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;os&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;join&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;BASE_PATH&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;cache&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;

&lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;webdriver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;Firefox&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;executable_path&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;DRIVER_PATH&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;uri&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sleep&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;5&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;quit&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;

&lt;span class=&quot;k&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&apos;Done&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h2 id=&quot;directivas&quot;&gt;Directivas&lt;/h2&gt;

&lt;p&gt;Como ha mostrado el ejemplo anterior, Selenium permite controlar el navegador a través del objeto &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;driver&lt;/code&gt; con directivas como &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;driver.get()&lt;/code&gt; para navegar a una URI.&lt;/p&gt;

&lt;p&gt;Las directivas que usaremos en este script son las cuatro siguientes:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;get()&lt;/code&gt;: Accede a una URI&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;back()&lt;/code&gt;: Retrocede a la página anterior&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;find_elements_by_css_selector()&lt;/code&gt;: Selecciona un elemento de la página&lt;/li&gt;
  &lt;li&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;quit()&lt;/code&gt;: Cierra el navegador&lt;/li&gt;
&lt;/ul&gt;

&lt;h3 id=&quot;guardado-local-de-cada-página&quot;&gt;Guardado local de cada página&lt;/h3&gt;

&lt;p&gt;Recordamos que el objetivo de este script es guardar una copia local de cada página que encuentre para posteriormente ser analizada con otras herramientas.&lt;/p&gt;

&lt;p&gt;Cuando el WebDriver carga una página, podemos acceder al contenido &lt;em&gt;crudo&lt;/em&gt; de la misma mediante &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;driver.page_source&lt;/code&gt;. De este modo, podemos guardar copias de las páginas HTML añadiendo:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;with&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;open&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;local_path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;s&quot;&gt;&apos;w&apos;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;k&quot;&gt;as&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;local_page&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;local_page&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;write&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;page_source&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;selección-de-elementos-html&quot;&gt;Selección de elementos HTML&lt;/h3&gt;

&lt;p&gt;De entre las muchas formas que ofrece Selenium de seleccionar elementos, el selector CSS es una de las más sencillas de utilizar. Este método será también utilizado en posteriores etapas con otras herramientas.&lt;/p&gt;

&lt;p&gt;Usando la sintaxis propia de CSS, podemos aplicar criterios de selección del mismo modo que lo haríamos en la hoja de estilos:&lt;/p&gt;

&lt;pre&gt;&lt;code class=&quot;language-Python&quot;&gt;links = driver.find_elements_by_css_selector(&quot;a.icon-arrow-right-after&quot;)
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;La llamada anterior devuelve una lista con todos los elementos de tipo &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;&amp;lt;a&amp;gt;&lt;/code&gt; (enlaces HTML) que además tengan la clase &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;icon-arrow-right-after&lt;/code&gt;. El nombre de dicha clase ha sido obtenido manualmente utilizando el inspector de código de Firefox. Una vez se conoce el &lt;em&gt;selector&lt;/em&gt; que se desea utilizar, su aplicación es automática por parte de Selenium.&lt;/p&gt;

&lt;h3 id=&quot;estructura-de-los-elementos&quot;&gt;Estructura de los elementos&lt;/h3&gt;

&lt;p&gt;Los elementos que devuelven las llamadas a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;find_*()&lt;/code&gt; son objetos internos de Selenium que contienen, entre otras cosas, los atributos del elemento HTML original. De este modo, el siguiente código permite obtener la URL a la que apuntaba dicho enlace:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;link&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;links&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;].&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;get_attribute&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;href&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;navegación&quot;&gt;Navegación&lt;/h3&gt;

&lt;p&gt;Con dicho enlace, ya podemos llamar a la función &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;get()&lt;/code&gt; para navegar a la siguiente página.&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;link&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;De un modo análogo se pueden obtener los enlaces a los pisos individuales y visitarlos uno a uno:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;n&quot;&gt;links&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;find_elements_by_css_selector&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;a.item-link&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;link&lt;/span&gt; &lt;span class=&quot;ow&quot;&gt;in&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;links&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;get&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;link&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
	&lt;span class=&quot;c1&quot;&gt;# Do something in the new page
&lt;/span&gt;	&lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;back&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;h3 id=&quot;check-captcha&quot;&gt;Check captcha&lt;/h3&gt;

&lt;p&gt;Una de las medidas que, sin duda, uno se puede encontrar cuando trabaja contra servicios protegidos contra &lt;em&gt;bots&lt;/em&gt; son los captchas: retos sencillos de resolver por un humano pero difíciles para un autómata.&lt;/p&gt;

&lt;p&gt;Para hacer frente a ellos, se ha optado por un enfoque semi-superviado en el que se observe el proceso ejecutarse de forma automática hasta que el portal web requiera la resolución de un captcha. En ese momento el script se detendrá a la espera de que el usuario supervisor resuelva el reto y éste pueda proseguir.&lt;/p&gt;

&lt;p&gt;Para logar esto de forma cómoda y automática programaremos que, ante la detección de un captcha, el script entre en pausa un tiempo determinado y, tras ese tiempo, compruebe si el captcha ya ha desaparecido:&lt;/p&gt;

&lt;div class=&quot;language-python highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;k&quot;&gt;def&lt;/span&gt; &lt;span class=&quot;nf&quot;&gt;check_captcha&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;():&lt;/span&gt;
	&lt;span class=&quot;n&quot;&gt;captcha&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;find_elements_by_css_selector&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;div.g-recaptcha&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
		&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;captcha&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
				&lt;span class=&quot;k&quot;&gt;while&lt;/span&gt; &lt;span class=&quot;nb&quot;&gt;len&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;captcha&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&quot;mi&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;:&lt;/span&gt;
						&lt;span class=&quot;n&quot;&gt;time&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sleep&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;mi&quot;&gt;5&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
						&lt;span class=&quot;n&quot;&gt;captcha&lt;/span&gt; &lt;span class=&quot;o&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;n&quot;&gt;driver&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;find_elements_by_css_selector&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s&quot;&gt;&quot;div.g-recaptcha&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Deberemos llamar a esta función cada vez que pidamos una página nueva con &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;get()&lt;/code&gt;. De este modo, cuando se vuelva de la función podremos estar seguros de que la página ya no contiene ningún captcha.&lt;/p&gt;

&lt;h2 id=&quot;algoritmo-completo&quot;&gt;Algoritmo completo&lt;/h2&gt;

&lt;p&gt;El algoritmo completo entonces resultaría:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Visitar la primera página del listado de items&lt;/li&gt;
  &lt;li&gt;Guardar una copia de dicha página&lt;/li&gt;
  &lt;li&gt;Obtener los enlaces a cada item&lt;/li&gt;
  &lt;li&gt;Para cada item
    &lt;ol&gt;
      &lt;li&gt;Visitar el item&lt;/li&gt;
      &lt;li&gt;Guardar una copia de la página del item&lt;/li&gt;
      &lt;li&gt;Volver atrás&lt;/li&gt;
    &lt;/ol&gt;
  &lt;/li&gt;
  &lt;li&gt;Obtener el enlace a la siguiente página del listado&lt;/li&gt;
  &lt;li&gt;Volver al primer punto&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;El código completo no ha sido compartido debido a que, en el proceso de desarrollo se implementó una estructura de guardado selectivo de páginas para facilitar las sucesivas ejecuciones del experimento que resulta totalmente innecesaria una vez de tiene controlado el proceso.&lt;/p&gt;

&lt;p&gt;En el caso de contar con tiempo suficiente como para realizar una versión más concisa y legible, está será compartida junto con el resto del proyecto.&lt;/p&gt;
</description>
        <pubDate>Fri, 08 Nov 2019 00:00:00 +0000</pubDate>
        <link>https://suresrm.github.io//2019/11/08/proteccion/</link>
        <guid isPermaLink="true">https://suresrm.github.io//2019/11/08/proteccion/</guid>
      </item>
    
      <item>
        <title>Web Scraping for Data Analysis in Python: a case study</title>
        <description>&lt;p&gt;En esta serie de artículos se cubrirá el proceso completo de Web Scraping y Enriquecimiento para el análisis de datos.&lt;/p&gt;

&lt;p&gt;El factor de diferencial que busca este trabajo es profundizar un paso más allá de los habituales “tutoriales en condiciones perfectas”, así como dotarlo de un escenario realista gracias al caso de estudio.&lt;/p&gt;

&lt;h2 id=&quot;estructura-de-los-artículos&quot;&gt;Estructura de los artículos&lt;/h2&gt;

&lt;p&gt;El proyecto constará de una serie de 5 artículos:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;a href=&quot;/2019/11/01/introduccion/&quot;&gt;Introducción&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;/2019/11/08/proteccion/&quot;&gt;Protección contra crawlers&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;/2019/11/15/scrapping/&quot;&gt;Scraping de datos&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;/2019/11/22/geoencoding/&quot;&gt;Geoencoding&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href=&quot;/2019/11/30/enriquecimiento/&quot;&gt;Enriquecimiento&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2 id=&quot;el-caso-de-estudio&quot;&gt;El caso de estudio&lt;/h2&gt;

&lt;blockquote&gt;
  &lt;p&gt;Estamos a punto de mudarnos a una nueva cuidad y queremos toda la información posible para decidir qué piso alquilar. Para ello buscamos crear una base de datos con todos los pisos de alquiler disponibles en esa cuidad, y que ésta contenga más información que la que el buscador de la agencia de alquileres nos ofrece.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Esta información:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Puede estar presente en la web, pero no ser “consultable” en su sistema de búsqueda:&lt;/strong&gt; el tipo de calefacción, certificación energética, el tipo de fianza, etc.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Puede estar implícita, pero no explícita:&lt;/strong&gt; el precio por metro cuadrado, dividiendo dos campos explícitos; la localización espacial del piso, infiriéndola a partir de su dirección, etc.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Puede no estar disponible en la web, pero podemos obtenerla cruzando datos con la localización espacial:&lt;/strong&gt; el tipo de zona en función del nivel de la renta o la edad media, la distancia a puntos de interés, etc.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2 id=&quot;fuente-de-datos-primaria&quot;&gt;Fuente de datos primaria&lt;/h2&gt;

&lt;p&gt;Se ha escogido el portal de datos de el Idealista como fuente principal de donde extraer los datos debido a dos motivos: gran cantidad de ofertas en la ciudad de estudio, Zaragoza; y fuertes medidas de seguridad contra crawlers y bots, que pretenden impedir la recolección de información de su portal (y hará este proyecto mucho más interesante).&lt;/p&gt;

&lt;p&gt;Por motivos de precaución ante posibles infracciones de términos y condiciones de uso, ni el código completo ni los datos extraídos serán públicamente compartidos.&lt;/p&gt;
</description>
        <pubDate>Fri, 01 Nov 2019 00:00:00 +0000</pubDate>
        <link>https://suresrm.github.io//2019/11/01/introduccion/</link>
        <guid isPermaLink="true">https://suresrm.github.io//2019/11/01/introduccion/</guid>
      </item>
    
  </channel>
</rss>
