miércoles, enero 23, 2013

Problemas en PHP

Tenía en mi colección de enlaces uno sobre los problemas que tiene PHP. Tras haber recibido otro enlace interesante comparto ambos aquí:

  1. PHP: a fractal of bad design
  2. PHP Sadness

lunes, enero 14, 2013

Escogiendo plataforma para una tienda online

Independientemente del motivo por el que llegues a la necesidad de tener que montar un tienda virtual, lo primero que tendrás que decidir es si vas a ir a por una solución SaaS o Hosted.

El usar un SaaS o no va a depender sobre todo de la funcionalidad que quieras ofrecer. Lo mas normal es que sirva un SaaS, como a mi.

Hosted

Te pueden llevar varios motivos a la decisión de no usar un SaaS. Yo diría que el principal es que lo que quieres implementar tenga unos flujos que no se van a adaptar a un producto ya enlatado.
Si optas por esta opción necesitarás o bien alguien técnico en el equipo  o tener algún contrato de mantenimiento para poder garantizar que la tienda esté online.

Posiblemente la decisión mas importante a tomar es en que lenguaje quieras trabajar: java, python, php, javascript, ...

En cualquier lenguaje vas a tener varias soluciones ecommerce [wikipedia](nada de python???), para poder escoger.
Si el lenguaje te da igual, la decisión se te complicará puesto que el abanico de productos se te abre. Es increible que incluso hoy en día siguen surgiendo productos nuevos... (no todo está en la wikipedia.. :p - ejemplo)

El siguiente paso será valorar cual de los diferentes productos ecommerce usar. Lo suyo será valorar cuales son tus necesidades e ir analizando producto por producto.
Una vez seleccionados los finalistas lo suyo sería hacer una prueba de concepto mínima, para ver con que se está mas cómodo (si hay tiempo).

Ojo! normalmente los CMS también tienen pluggins para convertirlos en una plataforma ecommerce.
Ver que opciones tenemos en los diferentes lenguajes daría para un post específico, que no de este.

Los machotes, podrán pasar de todo, y montarselo todo a pelusqui.

Ah! uno de los motivos por los cuales te podría interesar montarte tu mismo la tienda es porque quieras integrar alguna pasarela de pago distinga de la que ofrecen las plataformas. En otras palabras... que quieras integrar directamente alguna pasarela de pago de algún banco para abaratar costes. Esto lo recomendaría como una evolución a una tienda la operativa y justificándolo ya con datos reales.

SaaS

Si tienes claro que lo tuyo es un SaaS el siguiente paso es ver que opciones hay en el mercado.
Tras preguntar por aquí y por allá, ver que han usado conocidos, y buscar un poco por la red mis finalistas fueron:

  • Mangento go: el posible lider open source (como hosted) ya tiene solución SaaS.
  • Bigcartel: Un producto sencillo, ligado al mundo del arte (el de lo que estoy montado), y parece que bien valorado.
  • Shopify: Posiblemente el lider SaaS en cuanto a sencillez/funcionalidad.
  • Volusion: Una inclusión de última hora, por haber leído algunas reviews buenas sobre él.
Para decidir al igual que que si no usas un SaaS, hay que analizar las diferentes soluciones frente a los requisitos que uno tenga. Los requisitos pueden ser técnicos, funcionales, económicos, ... Es un tema ya personal.
Mis finalistas fueron Shopify y Bigcartel y al final que he quedao con Shopify. Acabo comentado lo que opino de cada uno... Podeis complementar la información buscando comparativas en la web...

Bigcartel

Fue el primero que probé. Inicialmente pensaba sólo comparar entre Magento y Shopify, pero me decidí a probarlo por su precio, que es realmente bueno. Un acierto.
Me sorprendió su sencillez. Me fuí encontrando dudas que el soporte me resolvió rápidamente y muy amablemente. Da buen rollo.
A nosotros nos servía porque cuadra perfectamente con el negocio a montar. Si el negocio es otro, en lugar de artistas habrá que trabajar con 'vendors', pero vamos... lo mismo da, ya que todo el css/html es modificable (un ejemplo de un site currado).
En principio y según reza en su pagina tienen un límite de 300 productos. Para nosotros ahora mismo no es un problema. Si algún día lo es dependerá del crecimiento. Sólo soporta Paypal como pasarela de pago lo que para nosotros tampoco es un problema ya que es lo que vamos a usar inicialmente. Aunque Shopify soporte mas pasarelas de pago hay que ver si en el pais de aplicación están disponibles...
Mi decisión en usar Shopify se basó en que la template basica de Shopify me gusta mucho más que la de bigcartel y ahora mismo no tengo tiempo de andar lidiando demasiado con el html/css/templates.
Otro detalle importante es que sólo puedes exportar las ordenes. No puedes exportar por ejemplo los productos. En Shopify puedes exportar los clientes, las ordenes y los productos.  

Shopify

Al igual que con Bigcartel es trivial poner el site a andar. El único pero que me encontré es que en la versión de prueba, no puedes probar pagos, con lo cual me quedé sin probar como es la generación de recibos. Al menos te muestran como son los pantallazos de la pantalla de ordenes cuando estas existan.
Aunque es mas caro que Bigcartel, si echas número el coste es razonable. Cuando ni siquiera has empezado es prematuro andar racaneando. Hay otros factores mucho mas importantes que ahorrar unos centimos por transacción. Ya habrá tiempo para eso...
Por lo demas... ningún pero. La funcionalidad que te da por defecto es muy buena y aparte puedes extender la funcionalidad con apps/pluggins. Hay muchas disponibles.
Tiene una fuerte comunidad alrededor y bastante documentación.
El coste inicial $29 + 2.0% por transacción está bien para empezar, y probar (coste de 5% aprox. contando paypal). Si las cosas van bien enseguida compensará migrar al siguiente escalón $59 + 1%.

Un detalle a tener que cuenta, que podría hacer que necesitéis una tarifa más alta de la que esperáis  Shopify cuanta cada variante de un producto, como producto. Si tenéis un producto con dos colores para Shopify serían dos productos. Si vais a manejar un gran número de productos con multiples variantes el número de productos se puede dispara rápidamente (imaginar una tienda de ropa con colores y tallas).

Magento Go

Sólo entrar en magento ya ves que el interfaz es mucho mas complicado y feo. Tiene mas opciones que el resto de los productos, pero para mi caso de uso innecesarias.
Las templates por defecto son mas féas.
La usabilidad también es peor. La creación de productos y categorías por defecto (puede que se pueda tunear para simplificar, no lo se...) es mas compleja.
Descartado por overkill y necesitar mas tiempo de diseño.
Su única ventaja actual para mi sobre shopify sería el precio.
Una ventaja futura sería la facilidad teorica de migrar de SaaS a Hosted.

Volusion

Volusion ni llegé a probarlo. Por un motivo, su precio para mi no era competitivo. Puede ser interesante en el caso de que requieras alguna funcionalidad que las otras plataformas no tengan; pero no era mi caso. 
No me gustó en que los comerciales me contactaron tanto por email como por teléfono.

I18n

Con respecto a los idiomas hay algunos peros: ninguno de los productos probados tiene el panel de administración en Español. Para mi no es un problema, pero para alguien podría serlo...

Y un pero bastante gordo para Bigcartel y Shopify. No soportan multilenguaje. Si queréis tener la tienda en mas de un idioma tendréis que montar mas de una tienda, con las implicaciones que tiene sobre todo a nivel de gestión (mas de un inventario, ...).

En este tema Magento Go, es el claro ganador.

viernes, diciembre 21, 2012

Websockets

A raiz de mi post anterior (Asyncronous web frameworks). Me ha dado por mirar el soporte a websockets en el mundillo python.

Node.js lo soporta con socket.io.
Play también tiene soporte para websockets (verdad que Play2 es muy interesante?).

En cuanto a python, hay varias opciones. Enlazando con el post anterior lo mas interesante sería combinarlo con gevent. Podemos usar tanto gevent-websocket como gevent-socketio.

Aquí tenemos otro post, en el que ademas usa zeromq.

Mas opciones pythoneras:


Por otro lado, la versión 1.3 de nginx incluirá websockets. Aunque parece que se puede hacer que lo soporte.

Referencias




Asyncronous web frameworks

Ahora mismo una de las guerras que tengo es progresar en javascript. Cosa que ya estoy haciendo sin necesitar empezar un nuevo proyecto, gracias a una migración que tenemos ahora entre manos en Youzee.

Dentro del mundo javascript, uno de los productos que están en ebullición es node.js. Y la pregunta que me hago es... ¿Me interesaría usarlo para futuros proyectos?

Estamos hablando de servir miles de peticiones concurrentes por segundo. Escalar puedes hacerlo con un webserver normal. Simplemente tienes que echarle dinero al tema y ya está. Se trata de poner mas hardware y a balancear. Pero porque vas a tirar dinero innecesariamente...

Por otro lado, tenemos el caso de uso del long polling, donde las conexiones son persistentes. Para este caso un servidor multithread no va a ser capaz de escalar con un gran número de clientes. Necesitaremos ir a un servidor basado en eventos y llamadas asíncronas  Ahí es donde node.js entra de lleno. node.js tiene la asincronía en su adn.

node.js claramente es una opción realmente interesante:

  • la comunidad de javascript está realmente activa y el lenguaje está cobrando mucha fuerza
  • node.js es realmente muy muy eficiente.
  • tienes la ventaja de poder compartir código con la parte cliente
  • nunca vas a tener el problema de bloqueos por llamar a código síncrono.

Ahora bien... ¿Que alternativas tenemos a node.js? A mi particularmente me interesa python. Es mi debilidad. Parece que la mejor opción sería usar gunicorn junto con gevent (pero no es la única, tenemos twisted, tornado, ...)

Creo que sólo una restricción me harían decantarme por gunicorn+gevent que el código tenga que ser python ya que lo mas eficiente a dia de hoy es node.js. Lo cual no quita a que el resto de servicios puedan ser python. Sólamente necesitaríamos que el frontend fuera node.js. 

Como ultima nota comentar que si no nos importa perder un poco de rendimiento gevent proporciona algo interesante y es que sin cambiar el modelo de programación consigues asincronía.

Otro lenguaje interesante es scala. Me apunté al curso de coursera, pero por falta de tiempo no puede realizarlo. Si podeis no dejeis de hacerlo que por lo visto es muy bueno.
Y en scala tenemos play2
Si hoy en día tuviera que usar algo corriendo sobre la jvm (restricción), lo tendría claro. Este sería el lenguaje. Se diferencia de javascript y python en que tiene tipado estático (y fuerte como python). Auna en él, la programación OO y la funcional.
¿La desventaja? Un lenguaje mas complejo de aprender que python y javascript. Pero play2 tiene una pinta fantástica.


Conclusión

play2, para mi ahora mismo tendría una curva de aprendizaje mas largo, por lo que para proyectos personales ahora mismo no lo consideraría. Eso si, gente que conozco que están empezando a trastear con play2 están encantados. 
A falta de ver mas comparativas a nivel de framework (por que a nivel de lenguaje, Scala tiene mejor rendimiento que python y javascript), profesionalmente lo consideraría si a nivel corporativo está como requisito que corra bajo jvm. 
Cuando me meta con Scala, fijo le doy una probazón.

A nivel general veo lo mas atractivo tanto para proyectos personales, como profesionales node.js. No sólo porque es el que da mas rendimiento, sino por que es ahora mismo lo está muy en boga. La opción de compartir código entre cliente y servidor es también muy interesante.

Python (gunicorm + gevent) lo usaría sin duda si el código tiene que ser python o sino quieres complicarte en crear servicios adicionales (nodejs + python services). La diferencia de rendimiento con node.js es pequeña. Vamos.. que es una combinación totalmente acertada.

Y no quería acabar sin comentar que se está trabajando sobre un estándar asíncrono para python 3.

Referencias

Javascript



Python



Scala









lunes, noviembre 19, 2012

Python + excel

La verdad... no se que fue primero si mi pensamiento en que estaría muy bien algo tipo excel en un entorno python o ver este artículo de Techcrunch, sobre DataNitro.

En Techcrunch referencian a:

Habrá que echar un ojo en profundidad para ver que ofrece DataNitro sobre ellos. Para usos comerciales ni Pyxll ni DataNitro son gratis, y parece muy similares.
Pyvot permite leer y manipular excel desde python, pero sin embargo no permite definir funciones en python que sean invocadas desde excel.

Echando un vistazo a este link del artículo,  me he encontrado con PyTools. Un plugging de Visual Estudio para python.

Buscando spreasheets en python sólo he encontrado:

Vamos... el único interesante parece Pyspread.

martes, noviembre 13, 2012

utc and localtime

Ayer vi una presentación muy interesante: Python for Humans.
En ella se comenta que a pesar de que python tiene su Zen (import this) hay modulos de la librería estándar que o bien no son simples o se podrían mejorar.

Un claro ejemplo es datetime. ¿Porque no se incluye la timezone por defecto?
En el momento que quieres hacer conversiones entre timezones, necesitas de modulos externos. Y de nuevo nos encontramos con otro de los problemas de python. La multitud de opciones... ¿Cual usar pytz o dateutil?

Os dejo con un programilla que convierte de utc a localtime usando los dos.


from datetime import datetime
import pytz
from dateutil import tz

# The problem with now and utcnow is that they don't set set tzinfo to None
now = datetime.now()
utcnow = datetime.utcnow()

# Wit this one, we are setting the tzinfo (using dateutil)
summertime = datetime(2012,7,1,23,30, tzinfo=tz.gettz('UTC'))

#--- Converson using dateutil
#    We are using naive timezones with the exception of summertime
#    Instead of using local datetime with tz.tzlocal() we are hardcoding it
print "*** Conversions using dateutil"
print 'now:', now.strftime('%Y-%m-%d %H:%M:%S')
print 'utc:', utcnow.strftime('%Y-%m-%d %H:%M:%S')
print 'summertime:', summertime.strftime('%Y-%m-%d %H:%M:%S')

UTC = tz.gettz('UTC')
HERE = tz.gettz('Europe/Madrid')
utcnow2 = utcnow.replace(tzinfo=UTC)
print 'utc localized:', utcnow2.astimezone(HERE).strftime('%Y-%m-%d %H:%M:%S')
print 'summertime localized:', summertime.astimezone(HERE).strftime('%Y-%m-%d %H:%M:%S')

#--- Converson using pytz
#    Pay attention to how we can create a utc datetime with timezone using now()
#    Instead of using pytz.utz we could also use pytz.timezone('UTC')
madrid = pytz.timezone('Europe/Madrid')
utc_with_timezone = datetime.now(tz=pytz.utc)
utc_withouth_timezone = datetime.utcnow()

print
print "*** Conversions using pytz"
print 'with:', utc_with_timezone.strftime('%Y-%m-%d %H:%M:%S')
print 'withouth:', utc_withouth_timezone.strftime('%Y-%m-%d %H:%M:%S')
print 'with localized:', utc_with_timezone.astimezone(madrid).strftime('%Y-%m-%d %H:%M:%S')

# astimezone cannot be applied to a naive time zone so we need to do like with dateutil
# and firtsly we need to localize the datetime
utc_withouth_timezone = pytz.utc.localize(utc_withouth_timezone)
print 'without localized:', utc_withouth_timezone.astimezone(madrid).strftime('%Y-%m-%d %H:%M:%S')
print 'summertime localized:', summertime.astimezone(madrid).strftime('%Y-%m-%d %H:%M:%S')

Salida:

*** Conversions using dateutil
now: 2012-11-13 12:17:01
utc: 2012-11-13 11:17:01
summertime: 2012-07-01 23:30:00
utc localized: 2012-11-13 12:17:01
summertime localized: 2012-07-02 01:30:00

*** Conversions using pytz
with: 2012-11-13 11:17:01
withouth: 2012-11-13 11:17:01
with localized: 2012-11-13 12:17:01
without localized: 2012-11-13 12:17:01
summertime localized: 2012-07-02 01:30:00

Vamos que da un poco igual usar uno que otro...
En el caso de pytz en lugar de usar localize() podríamos haber hecho uso de replace() como si se hace con dateutil.

Si teneis que usar offets, en lugar de zonas, dateutil los soporta. Mi balanza a priori se inclina para dateutil por ese motivo. Ejemplo:


from datetime import datetime
from dateutil import tz

HERE = tz.tzlocal()
OFFSET = tz.tzoffset(None, 3600)
UTC = tz.gettz('UTC')

now = datetime.now(UTC)
here = now.replace(tzinfo=HERE)
offset = now.replace(tzinfo=OFFSET)

print 'here:', here.astimezone(HERE).strftime('%Y-%m-%d %H:%M:%S')
print 'offset:', offset.astimezone(OFFSET).strftime('%Y-%m-%d %H:%M:%S')

jueves, octubre 25, 2012

py.test

Allá por el 2006 usé py.test. Me gustó su simplicidad. Muy pythoniano. ¿Lo malo? Que estaba vinculado a pypy.
Posteriormente he estado usando pyunit, que es mas crosslang.
Hoy he descucierto que ya no tiene la dependencia que tenía con pypy y que lo puedes usar más fácilmente: pytest.org.

martes, octubre 16, 2012

Mis editores de texto

Sublime Text 2 se ha convertido en mi editor de texto habitual (junto a vim).
Nunca he llegado a ser un crack del vim, pero me gusta. Desde que estoy en Youzee me manejo con un Windows 7, y tengo que confesar que ese probablemente sea el principal motivo. No llegue a tener un entorno de trabajo con vim para estar agusto al 100%.

Ahora tengo un entorno bastante decente, pero el que no le dedique otro intento tiene un culpable, que es ST2.

Hay alguna cosilla que puedo echar de menos con respecto a editores especializados, pero en esta vida no se puede tener todo...

Y el último editor que uso es  SciTE, en plan blog de notas y todolist.
No voy a justificar aqui el uso de ellos, simplemente quería comentar mis editores de trabajo y como los tengo tuneados.

Scite

use.tabs=0
tabsize=4
indent.size=4

font.base=$(font.monospace)
font.small=$(font.monospace)
font.comment=$(font.monospace)
font.text=$(font.monospace)
font.text.comment=$(font.monospace)
font.embedded.base=$(font.monospace)
font.embedded.comment=$(font.monospace)
font.vbs=$(font.monospace)


ST2


{
"color_scheme": "Packages/Color Scheme - Default/Monokai.tmTheme",
"default_line_ending": "unix",
"font_size": 10,
"theme": "Soda Dark.sublime-theme",
"translate_tabs_to_spaces": true,
"trim_trailing_white_space_on_save": true
}

Y uso los pluggings:

Package control
DocBlocktr
JsFormat
SublimeLiter

Vim

Vim lo tengo muy poco tuneado.
.vimrc


set expandtab
set shiftwidth=4
set softtabstop=4
set tabstop=4
set smartindent
syntax on


augroup filetypedetect
    au BufNewFile,BufRead *.pig set filetype=pig syntax=pig
augroup END

Y en .vim:


.
`-- syntax
    `-- pig.vim


Vamos... casi nada. Y de echo la sintaxis para pig, ya no la necesito :p





domingo, febrero 13, 2011

Youzee

En diciembre empezó una nueva etapa laboral.
Esta semana pasada lanzado lanzamos el teaser. Stay tuned!

viernes, diciembre 31, 2010

Desde XML a YAML

Estaba hoy revisando código de OpenERP, para revisar el comunicación cliente/servidor que tienen.
Es curioso que usan XML-RPC. No me lo había vuelto a topar desde que estuve con Fast Data Search.

He estado revisando el tema de los formatos en APIs, que tengo que hacer algunas APIs, ahí dejo unos enlaces recopilatorios:



Por cierto OpenERP en su nueva versión usan YAML.

miércoles, diciembre 15, 2010

PreNavidad 2010

En esta epoca de preparativos y compras ya navideñas me ha tocado pillarme unas pequeñas vacaciones. ¿El motivo? Mi etapa en Tuenti se acabó. Hoy era mi último día oficialmente.

A partir de mañana empiezo un nuevo reto, realmente interesante. Solo que voy a tener la oportunidad de estar casi en el equipo inicial de algo que esperemos pueda ver la luz. Que emocionante!
Y por el momento nada mas voy a decir, que es una spin-off. Tan spin-off de momento que ni siquiera tenemos oficina operativa al 100% :).

Cuando se me ofreció esta oportunidad no buscada, era dificil decir que no. Mas si cabe que los últimos meses en Tuenti estaban siendo un poco raros debido a cambios organizativos que me afectaron. Pero nada serio como para pensar en abandonar Tuenti.

Ya mirando para atrás, ha sido mas o menos un año y medio realmente intenso. De mucho aprendizaje (gracias chicos, sois unos cracks) y de un ritmo de trabajo fuerte (al final menos, todo hay que decirlo :P).

Cuando me incoporé en Julio del 2009 eramos unas 70 personas, contando a becarios (que en User Support eran unos cuantos) y actualmente se dobla ese número (y User Support está externalizado). Significa el pasar de conocer practicamente a todo el mundo a no. Vivir algo así en este periodo de crisis es increible.

En este tiempo la de cosas que he podido ver: Tuenti SMS, Tuenti Movil, Places, API, Games, 2 backend frameworks 1 frontend framework, rediseño total del look sitio, Groups, Pages, ...
Brutal!

A sido una etapa increible. Ahora de la mano de Telefonica (que no tiene nada que ver con mi salida) a ver como sigue evolucionando. Estaré al tanto... ;)

Va a quedar en la memoria un muy buen recuerdo:

  • esa cocina (menos mal que Teresa se encarga de tenerla en orden)
  • esa sala de juegos
  • esos madrugones para las releases (puagggggg, que poco me gusta madruga)
  • esos sofas
  • como al principio no me enteraba de lo que me quería transmitir el cherif
  • ese fundi y new fundi (after work cañas)
  • esas in office cerves
  • ese pedazo de equipo  (gente realmente involucrada, centrada y productiva)
  • mis dos monitores
  • una oficina con luz natural por doquier
  • estar al lado del concreso
  • tener tantos sitios alrededor donde poder comer
  • los hackmeups
  • ...
A ver como os las seguis apañando... No serán tiempos fáciles.   Facebook está desembarcando con todas sus armas. Nada mas anunciarse el OMV de Tuenti, Facebook ha anunciado ya su Places para España de la mano de 11870. Ufff, las espadas están en todo lo alto. Suerte chicos!

sábado, diciembre 11, 2010

¿Como seleccionar un SAI?

Necesito comprar un SAI, ya que estamos lanzando unos renders que llevan bastante tiempo y el suministro electrico no es muy estable.

La pregunta de turno, como seleccionar un SAI. La respuesta aquí.
En mi caso creo que me voy a ir a por un Riello IDialog ID80.
Lo fundamental es calcular cuantos watios vas a necesitar. Hay webs que te pueden hacer un calculo aproximado, pero vamos... Con la fuente de alimentación que tengas ya te vas a hacer una idea...

sábado, agosto 21, 2010

Un par de conceptos básicos de diseño software

Hoy he leído un par de articulillos sobre The Open-Closed Principle  y The Single Responsibility Principle.
Ambos de Robert Martin., y estraidos de objectmentor.com.

Sus correspondientes páginas de la wikipedia: OCP y SRP.

jueves, mayo 06, 2010

Privacy

Después de leer Top Ten Reasons You Should Quit Facebook me he ido a las condiciones de Tuenti. No las he leido, pero si he ido directamente a lo que me interesaba....

Al publicar contenidos en tu perfil -fotos, archivos, textos, vídeos, sonidos, dibujos, logos o cualquier otro material- conservas todos tus derechos sobre los mismos y otorgas a TUENTI una licencia limitada para reproducir y comunicar públicamente los mismos, para agregarles información y para transformarlos con el objeto de adaptarlos a las necesidades técnicas del Servicio.

Interesante diferencia ¿no?

miércoles, mayo 05, 2010

viernes, abril 23, 2010

PHP Performance

Estuve en la Campus Party EU de Madrid viendo una interesante charla de Rasmus Lerdord sobre PHP Performance.
Muy recomendable echarle un vistazo. Ver la charla es superrecomendable si teneis la ocasión aunque no esteis desarrollando en PHP.

Adjunto una entrevista a Rasmus, en la cual habla de HipHop.

domingo, abril 18, 2010

Evolving a Backend framework

Segundo artículo publicado en Tuenti: Evolving a Backend Framework.

Adjunto el contenido:

The duties of a Backend Software Architect at Tuenti include the maintenance and evolution of the Backend Framework. In this article we will talk about Tuenti’s framework evolution, share its pros and cons and briefly introduce its features without entering into many technical or architectural details (as they will be covered in future articles).

Historical Review

The software that runs www.tuenti.com changes continuously with at least two code deployments per week. The scope of these releases vary, but usually we release a lot of small changes that touch many different parts of the system. Of course sometimes our projects are really big and their releases get divided and released in series to reduce overall complexity and minimize risk.

Identical approach is appllied to framework releases. Currently the modifications are mainly subtle, but introduction of a framework had to be divided into few phases with some of them also decomposed into smaller ones.

The original version
Since its creation, the site runs over a lighthttpd, mysql and PHP. From the first version, no third-party frameworks have been used and all the software has been developed in-house (for the good or bad).

The first version of the “lib” was quite primitive from an architectural point of view, since as a start-up the primary aim of Tuenti was to reach the public fast and then evolve once the product was proven successful.

The transitional version
The transitional version was the one in place before we introduced the current framework. This code was using a framework built around the MVC pattern with a set of libraries supporting model definition and communication with storage devices (memcached and MySQL). At this point in time, the data partitioning was being introduced for both memcached and MySQL allowing Tuenti to scale much more effectively.

The use of memcached is very important for the performance of the site. When a feature was being implemented, the developer not only had to consider how the data is going to be partitioned in the database, he also had to decide what data is going to be cached in memcache, how the cache would work, and make sure that all interdependencies for data consistency are satisfied. The caching layer contained not only simple data structures, but also indexes, paging structures, etc.

The current version
Currently, newly developed domain modules use the new backend framework (that exchanged the old model and supporting classes) and we are gradually migrating modules from the transitional framework to the new one.

We have also designed and developed a new front-end architecture which is still under evaluation and testing. In the following months we will be posting more information about the framework and implemented solutions, so please be patient.

Some of the most important advantages of the new framework are:

  • standardization of data containers,
  • transactional access to the storage (even for devices not supporting transactions),
  • complete abstraction of the data storage layer.

In addition to above, the framework is introducing several concepts, among which you’ll find:

  • domain-driven development,
  • automatic handling and synchronization of 3 caching layers,
  • support for data migration, partitioning, replication,
  • automatic CRUD support for all domain entities,
  • object oriented access to data along with directly from containers (avoiding expensive instantiation of objects).

The framework is entirely coded in PHP and (so far) we have not moved any parts of the code into PHP extensions. This leaves us a lot of room for possible performance improvements but will reduce the flexibility of the code if we decide to make that step.

Selected framework features

A framework designed for a website like Tuenti has to address a lot of technical issues which you would not encounter in a standard website deployment. The problems arise on different fields: number of developers working on the project, scalability problems, the migration phases, and many more that appear as the site evolves over the time.

Although a deep explanation is out of scope in this article, let’s briefly see the mentioned features.

Transactional access to the storage
Systems using many storage devices require additional implementation effort to keep the data in a consistent state. We cannot completely avoid data inconsistencies (due to the delayed nature of some of the operations and failures), so we have to keep part of the consistency checks in the source code. Yet, we can minimize the impact and amount of problems in this area after implementing transaction handling within our application. This means that with more complex operations, that involve changes in several data sources we can keep a relatively high data consistency by implementing a design that defines “domain transactions” that relate to “storage transactions” assigned to different servers with different types of devices running on them.

This approach allows developers to focus on the logic and specific storage related cases, while the framework handles the transactions for most of standard operations automatically.

Complete abstraction of the data storage layer
A central point for the storage layer is a “storage target name”. These names are linked to several configuration data such as used storage devices, partitioning and/or replication schema, different authentication data, etc. In the domain layer, developers can write code focusing on logic and relations between domain entities and communicate with the storage layer as if it was one device (handling transactions as mentioned above).

This means that when there is a need to perform a data related operation, developers don’t need to worry about all the device specific details, caching, etc. Everything is handled automatically so (in the most common case) the data will come from memcache; if it was already used while handling this request – it will already be cached in the framework; or if the data has not been used for a while – it will come from MySQL since the cache has expired.

Standardizaton of data containers
Almost all data loaded into the system is stored in standard containers (DataContainer) that later are sub-classed to implement different logic for handling different types of data groups (Queue, Collection etc.). Implementation of standard containers allow us to integrate several features into the framework that not only speed-up development and reduce domain layer’s complexity, but also apply system-wide security and unify data access interfaces.

Drawbacks

Every architecture is designed with trade-offs in mind. This means that support of some of the architectural concerns is increased and for some decreased. In this case we have observed a higher memory consumption, bigger challenge in implementing particular performance related optimizations, and reduced flexibility of ways in which code can be implemented.

Currently developers have less freedom than when using the first version Tuenti’s back-end framework. Previously a developer could just write any SQL statement he wanted and decide whether to cache the data or not and how that caching should work to the last detail. There was more flexibility but the process was prone to errors and produced a lot of duplicated code (read: copy+paste or waste of time). We still need to provide a way for developers to write complex SQL queries that cannot be generated by the framework automatically but these are just exception as regular queries executed in Tuenti are very simple.

As was already mentioned, higher memory consumption and more challenging implementation of optimizations are drawbacks associated to the use of a more complex framework. Both CPU and memory consumption are not considered problems when we’re thinking about regular web requests. Standard response time was not affected in a noticable way, yet a glance at the back-office scripts execution statistics proves to us that there is still a lot of space for improvement in terms of memory usage and CPU consumption.

The root cause of higher memory consumption cannot be associated exclusively to the framework but also due to the fact that objects are cached in memory. Having a garbage collector is useless unless you release all references to objects. Caching is a very good solution to improve speed, but the code must provide ways to flush the cached data in order to make it usable in scripts that usually work on bigger amount of processed data then web requests do.

Evolution of the framework

A good framework design will allow for its evolution, but will define and enforce clear boundaries. Re-architecting the system is always a very difficult and expensive process, so one has to take into consideration all possible concerns (especially non-technical) and requirements defined for the system. It is also clear that the first version will never be the last one, so you need to be patient and listen to all of the feedback you’re receiving.

Once you have a stable version of your framework you need to convince the developers that it really solves their needs and that it will make their lives easier. Having your developers “on board” has several advantages:

  • they will suggest improvements and anything else that they feel that is awkward,
  • remove the communication barrier that will block your framework from “reality”,
  • speed up development process of the framework by streamlining ideas and effort.

When you are introducing a new framework, you also need to integrate it with the old one. This can be very hard and tricky. What you usually would like to do is to make the old framework use the new one. You need to maintain the old interface but run the new logic inside. Hopefully the old interfaces will make sense and you will not have to spend weeks trying to make “the magic” work in a technical world. You need to consider that the interface is not just the function signature and its arguments; you also have to respect the same error handling and influence of the old code on the environment.

As a framework developer you should never forget that the framework is there to help the people that are developing the functionality over it, however cool your framework is.


Memcache Migration

Tengo un par de artículos publicados en tuenti. El primero se titula Memcache Migration. Adjunto el contenido.

Tuenti is an example of one of the many sites that is currently using memcache to improve performance. It’s important to keep in mind that once you start using it your application performance will heavily depend on it. Consequently, if for any reason the memcache connectivity is lost, the user experience will not work as well as intended because it uses memcache as a key element that just failed.

When you deal with a lot of data and a lot of users, at some point you will need to replace a memcache server due to a hardware problem or a systems upgrade. When this happens you could simply disconnect it and replace it with another, but if you do so your cached data will be lost, which is normally something you don´t want. To avoid this problem, you need a mechanism to migrate the data from one server to another.

You might also decide at some point to change the format of the keys, or the format of the values. You could invalidate the keys, or empty the cache. But as previously mentioned, generally you want to make these changes while preserving the end user experience. The invalidation of keys can be done by including the version (or generation) in the key. By only changing the version, the previous key gets invalidated withouth any changes in memcache.

Another option for the migration is to stop the site and migrate the data during a maintenance window. If you have a small set of data this could be an option but you should avoid it. Users do not like maintenance windows.

Memcache Server Migration

In order to migrate a memcache server you need to design a mechanism to migrate the data from one server to another one while the site is up and running. The way to do this is by integrating the migration mechanism into your application code.

If you are facing this problem, it’s likely that you already have a framework that transparently accesses memcache, and redirects to data stored in the database when there is no data in memcache.

If a memcache server is accessed by an IP address, you need to change this inside your framework so it is accessed by your internal storage ID instead. That storage ID can have any information attached to it and one piece of information will be the IP address. When your application is coded to access the memcache via the storage ID you can configure it to transparently use any server. This comes in handy during memcache migration. In other words, you can get a memcache driver by its storage ID.

We could design a solution with the following objects.

The MemcacheFactory will return a memcache driver when its getObject($storageId) method is called. What driver is returned is transparent to client code.

If you manage a lot of data, you shouldn’t migrate all the keys that are being used at the same time. You should do it progressively by defining a migration rate. A migration rate of “0″ will mean that you don’t migrate any keys at all and a migration rate of 6 will trigger the migration of 6% of the keys. After determining the migration rate you need to decide the increment you want to apply to the rate and the frequency of the updates. The increment is directly related to the number of misses in memcache; it should be small enough to not increment them too much and avoid unhealthy spikes. The frequency is related to the number of accesses to the keys. It should be big enough to get the target keys used and consequently migrated. There are no magic values; the right values will depend on your data. For example, you could decide to increment 5% of the rate every 12 hours.

The method getObject would be something like (in PHP):

public function getObject($storageId, $forceRecreate = FALSE) {     if ($forceRecreate or !isset($this->drivers[$storageId])) {                 ...         $driver = $this->getDriver(...);         if ($migrationData !== FALSE) {                 $driver = new MigratingMemcacheDriver($driver);         }          $this->drivers[$storageId] = $driver;     }     return $this->drivers[$storageId]; }

The getObject method just with an ID instantiates the driver with all the data it needs (ip, port, …).

This leads to the question of how you choose what keys to migrate. You need a hash function that distributes your keys uniformly. The hash function you choose should be tested against real keys (you can get a set of them easily by sniffing the network). In the linked article, there is a good set of hash functions.
Now that you already have everything you need, how will it all work? If you are using a memcache driver that is being migrated it will perform operations depending on whether the key of the operation is being migrated. If the key is not in the process of being migrated the operation will go against the old server. If the key is being migrated the operation will go to the new server.

Let’s see it with code… For example the add operation in the MemcacheMigrationDriver would be something like:

public function add($key, $value, $ttl = MemcacheDriver::DEFAULT_TTL) {     return $this->getDriver($key)->add($key, $value, $ttl); }  private function getDriver($key) {     if ($this->migrationHelper->toMigrate($key)) {         $result = $this->migrationDriver;     } else {         $result = $this->driver;     }     return $result; }

The MemcacheMigrationDriver maintains internally the references to the old driver and to the new driver. It is a proxy object. The decision of whether a key is going to be migrated is delegated to a helper object.

If there is a ‘memcache miss’ after you load the data from the DB the data will the cached. With this solution you will be moving data from one server to other through memcache misses.

In the described solution all the logic related to the migration is distributed between the factory and the MemcacheMigrationDriver.

Memcache Data Format Migration

This migration happens inside a server, when a change in the format of the key or the format of the value is needed.

In case you want to change the value and in situations when it does not affect many keys you can just invalidate them. If you are using a framework you should be able to do it with a version tag by modifying the version and leaving the old entries in memcache to expire.

It´s important to keep in mind, however, that when the change affects a lot of keys or values you need a mechanism to do it progressively as in the ‘memcache server migration’.
Similar to the ‘memcache server migration’, this migration can be managed by a framework but the framework can´t manage the change in the format, therefore additional work is required by the developer who is migrating the data.
The worst scenario here would be when you want to change all the keys or all the values.

Migration of a key

Here we are going to show a different solution than the one described in ‘memcache server migration’. Instead of using memcache misses we are going to update the new entry without accessing the DB in case the old key exists in memcache.

The change of the key format should be done progressively because the number of accesses to memcache is increased.

In order to reflect that a memcache format migration is active you need to modify the configuration of the framework so that when a memcache miss is reached in a reading (in a writing we need to do nothing) a callback to get the old key is performed and the old key is used. A successful memcache hit with the old key will automatically trigger the update of the memcache.

To sum things up, during a migration of keys there are at least a couple of accesses to memcache in case of a miss. There are three accesses if the key is being migrated.

The implementation of this migration is done in a different layer than the ’server migration’. It could be implemented inside the driver, but it is generally preferable to implement it in a more external layer (in which you can have other functionality, like transactions for example).
The logic works only with one memcache driver. The code for a migration of keys would be similar to the code for a migration of values.

Migration of a value

As in ‘migration of a key’, you need a callback to get the right value from the key with the old format after it is received from memcache. The detection of the old or new format can be performed in another callback.

After the callback returns a new value the memcache is updated with it, so there are a couple of accesses instead of one to memcache.

Let’s see it with some code. The object that has made the change in the format has to implement the following interface.

interface MemcacheValueMigration {     public function getNewValue($key, $oldValue);     public function isOldValue($key,$value); }

So we could write the following code (as you can notice, the logic is outside the memcache driver):

private function getMigratedValue($key,$value,$object) {     $result = $object->getNewValue($key,$value);     if ($result !== FALSE) {         ...         $driver->store($key, $result);     }     return $result; }  private function recordFromMemcache($key,$object) {     ...     if ($this->migratingMemcache($object) and $object->isOldValue($key,$value)) {         $result = $this->getMigratedValue($key, $value,$object);     } else {         ...     }     ... }

Aditional considerations

Before implementing a migration mechanism consider whether it is truly necessary, because proper development requires some time (it won´t be implemented in an afternoon…) and a lot of testing. You have to be very sure it works properly before trying it out with real data.

The most dangerous migration is the migration of values because it can break down the system if the values get corrupted. You can provide mechanisms to rollback the operation or to invalidate the migrated keys, but they will have additional performance costs, as the access will be redirected to the DB for each invalidated key.

If something goes wrong in the migration of keys, you can just rollback to the previous format and all the new keys will be automatically invalidated. But its important not to forget that the values associated with the old keys could already be invalid.

The good news is that if there are any problem they will show up very soon.

Implementing the procedure incorrectly isn’t the only risk. Accesses to memcache shouldn’t be increased unnecessarily.

Conclusion

As it has been described in this article it is possible to implement a general mechanism to migrate memcache data within a framework.

Migration from one server to another can be done automatically, being activated by configuration.

Migration of keys or values inside a server needs some coding outside the framework. Helper classes should be provided by the framework.
Migration of a bit set of data should be done progressively.